Unlocking Safety Insights by Comparing Incident and Observation Text

Comparing text from two different corpora using Natural Language Processing (NLP) can provide valuable insights into safety incidents and observations within your organization.

3 min readData Science

The person asking this question has already done the hardest part: they suspect their observation program is performative, and they are looking for a way to prove it. That instinct matters more than any algorithm. Comparing incident text to observation text is not just a technical exercise. It is a direct test of whether your safety system is actually learning from what happens on the ground, or just checking boxes. If your observations are not describing the same activities, hazards, and behaviors that show up in your incidents, then you have not identified a correlation problem. You have identified a purpose problem.

What this means for you is that you do not need to become an NLP expert to get a defensible answer. You need to reframe the comparison. Topic modeling with LDA is a reasonable start, but comparing topic distributions across two separate corpora is misleading because the models are built independently. The topics will not align cleanly, and you will spend your time squinting at overlapping word lists instead of answering the actual question. A more direct approach is to train a classifier that tries to distinguish incident text from observation text. If the classifier can tell them apart with high accuracy, that is your evidence that the two sets of descriptions are talking about different things. If it cannot, then the texts are more similar than you think, and your observations may indeed be capturing the right risks. This is a simple, interpretable test that does not require deep fluency in topic modeling, and it gives you a number you can bring to leadership.

The practical next step is to prepare your text, strip out stop words, and run a logistic regression or a simple support vector machine with cross-validation. Report the accuracy and the most predictive terms. Those terms will show you exactly what language separates incidents from observations. If the top predictive words for incidents are things like "fell," "struck," or "caught," while observations cluster around "walked," "inspected," or "checked," you have your case. You will be able to say, plainly, that your observations are not describing the conditions that lead to incidents. That is not a vague feeling. That is a finding.

You do not need to solve every NLP challenge to make a strong argument. You need one clear signal that your observation program is misaligned, and the classifier approach gives you that. Stop trying to force a link between the two datasets. Instead, show that the language itself diverges. That divergence is the story. Bring that to your organization, and you will have moved from "I think this is busy work" to "Here is the gap, and here is where we need to focus." That is the kind of evidence that changes how a safety program operates.

From Data Science

I am not well versed in NLP, so hopefully someone can help me out here. I am looking at safety incidents for my organization. I want to compare the text of incident reports and observations to investigate if our observations are deterring incidents.

Read the original at Data Science