Human writers fail visibly: typos, missed deadlines, obvious first drafts. Models fail invisibly: fluent paragraphs containing one wrong number, confident claims with no source behind them, prose that reads beautifully and says exactly what the ten ranking pages already said. QA built for human failure modes waves the invisible failures through - which is how teams publish embarrassments at scale and conclude "AI content doesn't work". This article is the gate rebuilt for the actual failure modes: three checkable layers plus the one pass that stays irreducibly human. It is the quality chapter of content operations, and the gate our own daily article engine runs behind.
The inverted failure modes
The danger of model drafts is precisely their surface competence. A human editor skims for the signals that predicted problems in human work - clumsy prose, structural chaos, unfinished sections - and model drafts exhibit none of them while potentially containing: a statistic remembered wrong (fluently), a claim about a product that was true two years ago (confidently), a "study shows" citation to a study that does not exist (plausibly), and an argument identical to the SERP consensus (beautifully). The corrective principle: review the claims, not the prose. Prose quality is now table stakes; claim quality is where drafts live or die - which is why the first two layers below are about provenance and differentiation, not writing.
Layer one: factual verification
The layer's rule is inherited from the mission discipline: claims trace to sources, numbers trace to computation, and anything that cannot is cut rather than shipped. In a well-built pipeline most of this is enforced at drafting time - the writer model works from tool outputs and cites them, statistics come from the sandbox - which converts QA's job from "fact-check everything" to "audit the citation discipline": follow each sourced claim to its source and confirm it says what the draft says. Sampling depth scales with stakes: spot-checks for routine posts, complete verification for anything with legal, medical or financial surface. The failure this layer exists to catch is subtle - not fabrication usually, but drift: a source that says "up to 30%" becoming "30%" in the draft, three hops of paraphrase from true to false.
Layer two: sameness screening
Sameness is the failure nobody's checklist catches because every individual sentence is fine. The screen: put the draft against the content brief's falsifiable angle and against the current SERP's coverage, and ask two questions. Was the angle delivered - can a reviewer point to the sections that prove the claim the brief made? And does the piece contain anything - a finding, a dataset, a synthesis, a position - that the ranking pages lack? A draft failing both is the eleventh restatement, and no amount of polish rescues it: it goes back with the gap named, or the brief gets re-examined for whether the angle was ever achievable. This layer is why brief QA upstream matters so much - sameness caught at the brief costs a planning cycle; caught here it costs a production run; caught by the market it costs the library's reputation.
Layer three: compliance
| Check | Examples | Automatable? |
|---|---|---|
| Claims policy | Product capabilities stated within the approved envelope; no invented customers or numbers | Largely - against a maintained claims list |
| Voice rules | House style: punctuation rules, tone bounds, banned constructions | Yes - lintable |
| Legal surface | Comparative claims about named competitors sourced and dated; regulated-topic handling | Flaggable; judgment on flags stays human |
| Structural spec | The brief's required formats present: answer-first sections, tables, FAQ, links, sources | Fully - it is a schema check |
This layer automates almost completely - our own engine runs a structural lint with a scored checklist before anything publishes - which is exactly why it should never be the whole gate: automating layer three while skipping layers one and two produces perfectly formatted, on-brand, subtly wrong sameness.
The taste pass
Operating the gate at volume
At three posts a day, the gate survives only as a system. The mechanics: layers run in order of cost - structural lint first (instant), compliance second (automated), sameness third (a read against brief and SERP), factual audit sampled per stakes - with any layer's failure bouncing the draft before the more expensive layers spend attention. The reviewer works a queue of already-linted drafts, per the queue pattern. And the loop closes backward: failures tallied by type feed the pipeline - recurring factual drift tightens the drafting rules, recurring sameness tightens brief QA, recurring voice violations patch the lint - so the gate's pass rate rises over months, per the evidence-quality metric that tracks it. A gate whose failure types never change is a gate nobody is reading.
Frequently asked questions
What should QA check in AI-drafted content?
Three layers: factual verification (claims to sources, numbers to computation), sameness screening (was the brief’s angle delivered; does the piece add anything the SERP lacks), and compliance (claims policy, voice rules, structure) - plus a human taste pass.
Why is skimming not enough for model drafts?
Model failures are fluent and confident: wrong numbers in clean sentences, plausible citations to nothing, beautiful restatements of what already ranks. Prose signals no longer predict problems; claim provenance does.
What is sameness screening?
Checking the draft against the brief’s falsifiable angle and the current SERP: point to the sections that deliver the angle, and name what the piece contains that ranking pages lack. Failing both means rework, however polished the prose.
How much of the gate can be automated?
Structure and compliance almost fully, factual verification partially (citation-presence checks automate; source-agreement audits sample), sameness partially (similarity flags help; the angle judgment is editorial). Taste automates not at all.
How does the gate improve over time?
By feeding failures backward: tallied failure types patch the drafting rules, brief QA and the lint. Rising pass rates with stable standards is the health signal; a gate with unchanging failure types is not being read.
Sources
- Google Search Central - people-first content guidance (the standard the gate serves)
- Anthropic - Building effective agents (grounding outputs in tool results)
Every playbook on this blog ships as a runnable mission.
Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.