Research · Aug 19, 2026 · Updated Sep 8, 2026 · 2 min read
Can an LLM Draft a Usable CAPA? We Measured It
We scored 240 AI-drafted corrective actions against what our QA team actually shipped. Where the drafts held up, where they failed, and why the failures cluster.
Every food safety software vendor now claims AI can draft your corrective actions. We wanted a number instead of a claim, so we ran one: 240 deviation records from our own certified facility, each drafted by the model, each scored against the version our QA team ultimately approved.
How we scored it
Each draft was graded on four axes: root cause plausibility, corrective action specificity, regulatory grounding, and whether the effectiveness check would actually detect recurrence. A draft passed only if a QA manager could sign it with edits taking under five minutes.
What we found
Roughly three quarters of drafts passed the five-minute bar on routine deviations: monitoring misses, documentation gaps, single-point equipment failures
Drafts failed hardest on root cause, defaulting to retraining when the honest answer was a design or scheduling problem
Effectiveness checks were the weakest section: the model proposes verification that confirms the action happened, not that it worked
Regulatory citations were accurate when the source requirement was in context, and confidently wrong when it was not
The "retrain the operator" reflex
The most instructive failure mode: when the record did not contain enough to find a root cause, the model filled the gap with the industry's own worst habit, blaming the person and prescribing training. It learned that from us. The fix is structural: the draft has to be grounded in the deviation record, the equipment history, and the schedule, and when the ground is thin, it should say so instead of guessing.
What this means for review
The human is not there to fix grammar. The reviewer's real job is the two sections the model is worst at: interrogating the root cause and demanding an effectiveness check that would actually catch recurrence. A review process that knows where the failures cluster reviews five times faster than one that reads every word with equal suspicion.
Run a food safety program? See how Beacon keeps it audit-ready.
See Beacon in 20 minutes