The short version: AI writing tools are good at a fast first draft, plain language, and consistency across a stack of documents. They are not good at knowing your site. A large language model can produce a requirement that sounds exactly right and is quietly wrong, because it predicts likely wording rather than checking reality. Use it to draft. Keep a competent human as the owner who verifies every line against the actual equipment, task, and crew.
Ask a chatbot for a toolbox talk on ladder safety and you will have a clean, readable page before your coffee is cool. Ask it for the skeleton of a lockout policy and it will give you one. That speed is real, and pretending otherwise helps nobody.
The useful question is not whether to use these tools. Plenty of practitioners already do. It is where they help and where they quietly hurt, so you can keep the first and design out the second.
Where it genuinely helps
The strongest use is the blank page. Staring at nothing is the slowest part of writing a JHA or a program, and a model will fill that page with a reasonable starting structure in seconds. You are no longer writing from zero. You are editing, which is faster and easier to do well.
Plain language is the second win. Feed it a stiff, clause-heavy paragraph and ask for a version a new hire can read, and it usually delivers. Safety documents that nobody understands do not protect anyone, so readability is not a nicety. It is the point.
Consistency across a stack of documents is the third. If you have forty toolbox talks written by six people over ten years, a model is good at matching them to one format, one heading style, one reading level. It will catch the JHA that calls it a "harness" while the next one says "fall arrest system" and flag the mismatch.
Translation is the quiet standout. A crew that reads the safe work method in its own first language understands it better than one squinting through a second. A model gives you a fast working draft in Spanish, Vietnamese, or Tagalog for a bilingual supervisor to check before it reaches the floor. That is a real gain for a real crew.
Where it quietly hurts
Here is the failure that matters. A language model generates text by predicting likely wording, not by checking facts. So it will sometimes produce a requirement that reads as perfectly authoritative and is simply invented. Researchers call these hallucinations, and a 2024 study presented at NAACL found that models produce them even about facts they demonstrably contain, with no visible signal to the reader that anything is off.
You do not have to imagine the consequences. In Mata v. Avianca, two attorneys filed a brief built on six court cases that ChatGPT had fabricated, complete with convincing citations and quotations. None of the cases existed. A federal judge sanctioned the lawyers in 2023. The tool did not fail loudly. It failed in a way that looked exactly like success.
Now move that into a safety document. A model might state a torque value, a clearance distance, an inspection interval, or an air-monitoring threshold that sounds standard and does not match your equipment, your chemical, or your task. It might invent a step in a lockout sequence that is plausible and wrong. On the page it reads no differently from the parts it got right, which is exactly what makes it dangerous.
The line that never moves
A competent human owns the document. That does not shift no matter how good the tool gets.
Owning it means one qualified person verifies every line against the real site: the machine that is actually on the floor, the chemical actually in the tank, the task the crew actually performs. It means the AI wrote a draft, and a person who could have written it unaided confirmed each requirement is true here. The model's own confidence that a value is correct is not verification. It is the thing you are checking.
OSHA's guidance on identifying hazards makes the same underlying point without mentioning software at all: the people closest to the work are your best source for what is actually risky about a task. A JHA is built by walking the job with the crew, step by step. A model has never walked your job. It can format the walk-through beautifully once you have done it. It cannot do the walking.
A practical way to run it
Treat the tool as a fast junior writer, not a subject expert. Let it draft, translate, plainen, and standardize. Never let its output reach the floor without a named reviewer who signs their name to it.
Keep the source of truth human. Values, thresholds, sequences, and site-specific facts get checked against manuals, standards, and the equipment itself, not against the model. Feed the model your real data and ask it to arrange the writing, rather than asking it to supply the facts.
Used that way, an AI assistant gives you back the hours you used to spend on blank pages and formatting, and puts them where they belong: on the walk-through, the crew conversation, and the verification only a person can do.



