'Does It Read as AI' Is the Wrong Test

Alexander Chua
8 min
'Does It Read as AI' Is the Wrong Test

The short version

  • An AI-detector score grades the surface of a page. It says nothing about whether the page holds anything a reader cannot get from the first result.
  • In the GEO study (Princeton, Georgia Tech, Allen Institute for AI and IIT Delhi, KDD 2024), adding statistics lifted AI-answer visibility 40%. Fluency optimisation, what a de-AI rewrite buys, lifted it 15 to 30%.
  • Removing AI tells changes the register of a page and nothing else about it.
  • Put the mechanical rules in a script. Spend the review on the question a script cannot answer: what on this page came from us.

The question worth asking about a page is what it contains that a reader cannot get from the first result. We graded drafts on how human they read, and that turned out to be the wrong quality test. Ranking follows helpfulness and first-party evidence, and a detector score measures neither.

An AI-written batch of twenty posts missed our own documented content playbook entirely.

What does ‘reads as AI’ actually measure?

Surface features, and only those. Em-dash density, the sandwich where a claim gets restated after its own bullets, bullets that all open the same way, paragraphs of an even length. Each is a real tell worth removing. None describes what the page contains.

What that batch missed was a documented standard rather than a register. A rewrite that removes tells does not put a standard back.

What moves visibility in AI answers, measured by tactic (GEO study, Princeton, Georgia Tech, Allen Institute for AI and IIT Delhi, KDD 2024):

TacticMeasured liftWhat it requires from you
Adding statistics+40%A number from your own work
Adding quotations+40%A real source, quoted
Citing sources+30 to 40%Publisher and year inline
Fluency optimisation+15 to 30%Cleaner prose, which a rewrite produces
Schema markup2.5x higher chance of appearing in AI answersThe page declaring its type in JSON-LD

Why is fluency the weakest lever?

Three of those tactics need something a generator cannot supply on its own. A statistic comes from work somebody did, a quotation from a person who said it. A citation has to name a publisher that exists.

Fluency is the one a rewrite manufactures alone, and it carries the smallest percentage lift on that list.

What replaced the detector test here?

Two files and one question. The fact bank holds the numbers that may be published, tiered by how each one may be used. The experience bank holds a row per engagement worth writing from, with its lesson and the label that replaces the client’s name.

The question is what a reader gets here that the first result does not. A draft with no answer goes back, however it reads. Both files go into the brief, because a standard the model cannot see does not exist.

Where do the AI tells still matter?

In a script rather than in a review. Banned characters, banned constructions and missing fields are all checkable, and 44 content pages on one account sit behind a gate that checks them at publish.

Google ranks helpfulness and first-party evidence, not an AI-detector score.

Put a number from your own work on the page and name where it came from. Then say something a competent practitioner could disagree with.

The two failures either side of this one are what happens to numbers in a fluent draft and why a clean library still gets no citations.

Frequently asked questions

Does Google penalise AI-generated content?

Ranking follows helpfulness and evidence rather than authorship. What is measurable is which tactics lift visibility in AI answers: adding statistics gave 40% and citing sources 30 to 40% in the GEO study (Princeton, Georgia Tech, Allen Institute for AI and IIT Delhi, KDD 2024). Nothing there concerns who typed it.

How do you tell whether AI-generated content is good enough to publish?

Ask what is on the page that a reader could not get from the first result. A number from your own delivery data, or a first-hand account of the work. Fluent prose with no evidence fails that question however human it reads.

Are AI content detectors reliable enough to use as a quality gate?

No. A detector grades surface features, so a rewrite moves the score without changing what the page contains. A draft can pass while carrying nothing first-party, and the tactics that do lift AI-answer visibility are the ones a detector cannot see.

What is information gain in content?

Information gain is what a page adds that the results above it do not already carry: your own measured numbers, or a first-hand account of work you did. A position a competent practitioner could argue with counts too. It is the hardest thing for a generator to supply.

Should you still remove AI writing tells?

Yes, in a script rather than a review. Banned punctuation, banned constructions and missing required fields are mechanical, and mechanical rules belong in a publish gate. 44 content pages on one account sit behind that check.

What should a content review spend its time on?

The parts a script cannot judge: whether the argument holds, and whether every number resolves to a source. Once the mechanical checks run in the publish path, the review budget goes to substance.

Sources

Alexander Chua

Alexander Chua

Co-Founder, PipelineRoad. Building companies and observing the world across 40+ countries. Writing about company building, go-to-market, capital formation, and the lessons in between.

More about Alexander

Newsletter

Chua Network Letter

Occasional essays on company building, global observations, and clear thinking. No spam. No SEO bait.