The short version
- A generator states a number it has never seen in the same register it states one you supplied. Confidence tells you nothing.
- Keep one file with a row per publishable number: value, source, and whether it can be used. Nothing goes on a page without a row.
- Run the fact check as its own adversarial pass. A proofread reads for sense, and a fabricated figure makes perfect sense in its sentence.
- Hedges and scope belong to the number. Roughly 92% of one client's raw traffic being bot is a different claim from 92%, and different again from a portfolio figure.
Fluent output is confidently wrong. A generator states a number it has never seen in the same register it states one you supplied, and nothing in the sentence marks the difference. Every number in a draft gets checked against a source file before publish, in a pass whose only job is to break it.
The failure it prevents is a buyer asking where a figure came from and getting silence.
Why does a generated number arrive wrong?
A number is a token like any other. The model produces whichever figure fits the shape of the sentence, and a plausible one fits better than an awkward true one. Nothing separates a supplied number from an invented one.
Outright invention is the easy case. A hedge falls off, so roughly 92% becomes 92%. A figure measured on one account gets restated as an average across the book, or a third-party finding arrives phrased as ours.
Every number we publish resolves to one of these rows first:
| Tier | What it holds | Example | Can it lead a page? |
|---|---|---|---|
| A | Verified, ours, published before | 8 client accounts across 11 brands | Yes |
| B, delivery | Verified delivery data, first use here | 0.11% average organic CTR for a month, Search Console | Yes, source named |
| B, AI production | Measured in our own production runs | 3 of 3 ad images on-style with a reference image, 2 of 3 without | Yes |
| C | Blocked pending verification | Any revenue, margin or retainer figure | No |
| External | Third-party research | GEO study, KDD 2024 | Yes, publisher and year inline |
| Closed | Commercially sensitive | Contract values, renewal terms, capacity | Never |
What goes in a fact bank?
One row per publishable number, carrying the value, its source and where it has been used. The tiering does the work.
Contract values, renewal terms and delivery capacity are ours to know and not ours to publish, and the file records them as closed.
What is an adversarial fact-check pass?
A separate pass with one job: assume every number is wrong, then go and find the row. A proofread will not do it, because a fabricated figure makes perfect sense in the sentence it was built for.
The rules are narrow. The value matches the row rather than approximating it. The hedge travels with the number, and so does the scope: a month of one account’s Search Console data never becomes how our clients perform.
How do you handle third-party statistics?
Publisher and year inline, never in a footnote. The GEO study from Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi (KDD 2024) is cited that way wherever it appears, and never rounded or extended.
Two more rules earn their keep. A cited URL is checked to resolve before it ships. Anything that moves, such as keyword volume, carries its pull date. A volume figure with no date names no month.
A number with no row behind it does not get hedged into safety. The sentence gets rewritten without it.
What is the bar for a published number?
A call rather than a page. If a buyer asks where a figure came from, the answer arrives in five seconds or the figure should not have shipped.
The mechanical half runs as a script, the way every house rule that survives has to.
The two checks that run either side of this one are counting the built page instead of trusting the draft and the quality test that replaced our AI-tell review.
Frequently asked questions
Why do AI models produce fake statistics?
Because a number is a token like any other and the model produces whatever figure fits the shape of the sentence. A plausible figure fits better than an awkward true one, and nothing marks which is which.
What is a fact bank?
A single file holding one row per publishable number: the value, where it came from, and whether it can be used. Ours tiers rows into what can lead a page, what can be used with its source named, and what stays closed.
How do you fact-check AI-generated content?
As a separate adversarial pass, one number at a time, assuming each is wrong until a row in the source file matches it. Reading for sense does not work, because a fabricated figure reads well in the sentence it was written for.
What is the most common way a true number gets corrupted?
The hedge or the scope falls off. Roughly 92% becomes 92%, or a figure measured on one account for one month becomes a portfolio average. Both read as confidently as the original row and neither is the same claim.
How should third-party statistics be cited?
Publisher and year inline on the page, with the figure unrounded and unextended. Check the URL resolves before shipping, and date anything that moves, such as keyword volume, because a figure without a pull date names no month.
What do you do when a number has no source?
Rewrite the sentence without it. Softening a figure into a vaguer version of the same claim keeps the problem and hides it. The test is whether a buyer asking on a call gets an answer in five seconds.
Sources
- Chua Network delivery data across 8 client accounts (internal fact bank)
- Chua Network engagement records, anonymized (internal experience bank)
- GEO study, Princeton, Georgia Tech, Allen Institute for AI and IIT Delhi, KDD 2024