AEO Content AI: free AI content detector AEO Content AI: free AI content detector

What Makes AI-Drafted Content Read Human: First-Hand Data

Give the draft material only your company could publish, then edit the surface. On one site we measured, record-based sections read AI 37% of the time, against 71% for the rest.

A content writer reviewing printed business records and handwritten notes next to a laptop, integrating first-hand data into an article draft
Three things content teams believe. Myth or fact?
Call each one, then see how other readers called it.
1 Better prompts alone can make AI-drafted content pass detectors.
2 All sections of an AI-drafted article read machine-like at the same rate.
3 The sections most likely to read human are the ones a model did not write from scratch.

Quick Answer

How do you make AI content sound human?

Give the draft material only your company could publish, then edit the surface. On one site we measured, record-based sections read AI 37% of the time, against 71% for the rest.

Treat that as one strong result, not a rule: on a payments company's site, whose articles used general industry figures, the same comparison came out at 98% against 100%. Surface editing has limits of its own. Removing reused phrases, dash habits and fixed skeletons made our articles stop looking alike, yet 81-100% of every generated prose section type still read AI to a detector trained on that pipeline's output. The order this suggests, as guidance rather than a finding, is to collect the records before drafting, because no revision pass adds specificity that was never collected.

Did this answer your question?

What separates AI-drafted content that reads human from content that doesn't?

In our data, the dividing line was who supplied the material. Bios from real profiles, curated link lists and, on one site, record-based sections read AI less often than generic prose.

We scored 2,625 sections of 40 words or more from 146 AI-drafted articles, written for 13 websites between 1 and 13 September 2026, one section at a time. The detector was our own: a ModernBERT classifier plus a separate predictability check from a pair of open Qwen2.5 models, which splits text into passages of about 300 words. It calls a passage AI at a score of 0.998 or higher and human below 0.5, and leaves anything in between uncalled. It was trained on the same pipeline's output, so it recognises that output by construction.

As a control, we cut verified human Medium posts into 50-, 100-, 200- and 300-word chunks. The detector called 1.0%, 0%, 0% and 0% of them AI, so short sections were not flagged just for being short.

Section type Sections Read AI Read human
Decision framework16100%0%
Prediction or outlook chart9499%0%
Before and after comparison9994%2%
What will matter most (outlook)10289%5%
Opening paragraph14686%3%
Intro teaser14487%3%
Closing paragraph14287%5%
Myth vs fact13587%3%
Key takeaways10287%2%
FAQ14481%4%
Body section76081%7%
External resources list5242%21%
Call-to-action banner12536%33%
Author bio398%51%

What reads human in that table is mostly what the model did not write from scratch: bios built from a real profile, calls to action and curated resource lists. Fixed rhetorical shapes, such as decision frameworks, outlook charts and before-and-after comparisons, read the most machine-like. Reused phrasing was not the difference. Phrases shared across different websites made up 0.5% or less of every prose section, yet 81-100% of every generated prose section type still read AI.

Practitioner advice on humanizing AI writing mostly targets the surface. A Substack guide on making AI writing more human recommends turning Wikipedia's "Signs of AI writing" page into a reusable humanizer prompt, and argues that "The substance is fine most of the time. It's usually the packaging." Our section data does not settle that for every site, but in our pipeline, packaging fixes alone left most generated sections reading AI.

Why a company's own records help is a plausible explanation, not a finding: counts, dates and places that only one company holds are details a model cannot produce from general knowledge. The one site where we saw the effect clearly is the evidence. The payments company, whose articles used general industry figures and showed 98% against 100%, is the limit.

This article covers:

  • What makes generated sections read machine-like, and what can change before drafting
  • Where style habits such as dash rate explain part of the picture, and where they do not
  • What our detector and four open detectors show about our own pipeline's articles
  • Why humanized drafts still read AI to a detector trained on their pipeline

The substance test, described later in the section on humanizer passes, is a quick editorial check for generic passages. It is guidance, not a detector.

What actually makes AI content read machine-like, and can it be changed before drafting starts?

Fixed section shapes and generic material read the most machine-like in our data. The material, at least, is decided before drafting starts, so that part can change before a word is written.

A January 2026 Medium post on beating AI detectors documents one team's attempts. Their Claude pipeline, prompted with E-E-A-T principles, citation standards and formatting rules, produced output the author calls "excellent for SEO/GEO" but "terrible for humanness score" on GPTZero. A loop that regenerated whole drafts left the score at zero. Letting the model call GPTZero as a tool landed scores between 0 and 25. Rewriting only the flagged sentences, one at a time, worked best of the three. The author still says no method works consistently 90% of the time or more, and that "Tools like GPTZero, ZeroGPT, and Quillbot improve their models almost weekly, breaking what worked yesterday."

That test measured one commercial detector, not what makes text read human in general. Our own measurements point at structure and material. In 146 AI-drafted articles, 56 contained at least one passage that read human, but an article is judged by its most machine-like passage, so 144 of 146 were still called AI. The opening was the most machine-like part: 87% of first passages read AI, against 64-75% for later passages.

Some writers now change the first step instead of the last. Developer Elio Struyf describes in Getting interviewed by AI to write your technical content or blog posts how an AI interviewer, in sessions of about 10 minutes, drew out the material for three of his blog posts. His reason: "The writing process filters out a lot: the tangential insights, the 'oh by the way' details, the thought process behind technical decisions." He is careful to add that the approach does not automate writing ("it doesn't, not really").

This article separates what we measured from what we infer. Measured results are presented as findings; practical advice drawn from them is labelled as guidance.

How will the gap between human-reading and machine-reading content shift over the next 12-24 months?

We expect the gap to follow the material: sections built from real profiles, curated links or a company's own records should keep reading human more often than generic drafting. None of this is certain.

The starting point is measured. Author bios built from a real profile read human 51% of the time in our September 2026 scoring, calls to action 33% and curated resource lists 21%. Body sections, the bulk of every article, read human 7% of the time and AI 81% of the time. Those rates come from all 146 articles across 13 websites. The 37% against 71% result for record-heavy sections is a separate comparison, from one sports equipment manufacturer's site.

Why the model-light blocks read human is a plausible explanation rather than a finding: they carry text the model did not write from scratch. Whether that holds as models and detectors change is the forward-looking question. Our view, labelled as prediction, is below.

Signal Prediction (12-24 months) Weak signal now Why it matters
First-hand data as authenticity marker Material built on a company's own records will keep reading human more often than generic drafting on the same site. On one sports equipment manufacturer's site, record-heavy sections already read AI 37% of the time, against 71% for the rest. On a payments company's site that used general industry figures, the comparison was 98% against 100%. Buyers choosing writing partners should weigh access to original records alongside stylistic polish.
Platform-embedded drafting floods professional publishing As drafting tools are pitched to major publishers and low-cost AI-interview products reach independent creators, the volume of interchangeable machine prose will grow. In July 2023, Nieman Lab reported that Google had pitched its Genesis news-writing tool to The New York Times, The Washington Post and News Corp. AI-interview products such as ContentPod start at $9.99 per month. Readers and buyers will see more interchangeable prose, which should make verifiable originality easier to notice.
Cosmetic humanizer tools plateau Humanizers that change only surface patterns will lose ground as detectors are retrained on their output; first-hand material will prove the more durable investment. A classifier trained on our own pipeline's humanized output separates those articles perfectly (AUROC 1.000). Sun et al. (ICML 2025) told ChatGPT, Claude, Grok, Gemini and DeepSeek apart with 97.1% accuracy, and the result persisted after rewriting, translation and summarization. Anyone budgeting for humanizer subscriptions should expect results to vary by detector and plan for original substance as well.

The forecast changes if new models start producing specific, varied prose by default, which would shrink the premium on first-hand records. Our September 2026 data does not show that yet.

Does the style signature of AI drafting explain what makes content read machine-like?

Only partly. Style habits such as dash rate are real and measurable, but removing them in our pipeline did not stop generated sections reading AI, and some human writing gets misread too.

Human writing is the first check. On 2,868 Medium posts archived before the end of 2020, the previous version of our detector called 4 AI (0.14%). Three were institutional prose (a government programme update, a health policy analysis and a corporate design case study), and one was a list of 100 headline-style post titles. None were AI-written. The current version calls 2 of the same 2,868.

Our research reads that pattern this way: the human writing most often mistaken for AI is neutral, well-organised institutional prose, because that is the register AI models imitate by default. That is a plausible explanation, not something four posts prove. It fits a second test, where the previous version called 8 of 637 US federal agency documents AI, above our 0.5% cap, before retraining brought it down to 2.

Style drift in AI drafts is also measurable. In September 2026, a fleet of AI-drafted articles contained zero em-dash characters, because a sanitizer had swapped each one for a spaced hyphen, yet carried 12.9 dash-punctuation marks per 1,000 words, against 0.8 per 1,000 words in one company's own pre-AI blog. After we replaced the writing rule with a budget of 2 dash marks per 1,000 words, new articles measured 1.69 dash marks per 1,000 words. That brings the surface closer to the human baseline. It does not change what the article is built on.

Cleanup of that kind changed the surface, not the verdict. Removing reused phrases, dashes and fixed skeletons made our articles stop looking alike, but 81-100% of every generated prose section type still read AI to a detector trained on the pipeline, with shared phrases at 0.5% or less. Outside research points the same way: Reinhart et al. (PNAS, February 2025) describe a noun-heavy, information-dense model style, with a larger gap from human writing for instruction-tuned models than for base models.

Readers lean on style cues too, and those cues misfire. In a Stanford HAI study (Hancock et al., March 2023), people judging dating, professional and hospitality profiles told human from AI text with 50-52% accuracy, and wrongly read grammatical correctness, first-person pronouns, family references and informal language as signs of a human writer.

We therefore report two readings separately. Origin is who most likely wrote the text: AI, a person, or not called. Style is how generic or distinctive the writing reads. The two do not move together. On a sports equipment manufacturer's site, 288 posts by one human author and 30 AI-drafted buying guides scored in a similar style band (57 against 68 on our style scale), while our detector flagged 0.7% of the human posts and 73% of the guides.

So style explains part of what reads machine-like, and editing style is legitimate work in a production pipeline. It does not settle authorship on its own, and it cannot add material that no generic draft could have included.

What does the detector data show about which AI-assisted articles actually read human?

Four open detectors called none of 223 held-out articles from our pipeline AI, yet a classifier trained on that pipeline's output separated them perfectly. Reading human depends on who is asking.

The open detectors were desklib, Fakespot, Binoculars and Fast-DetectGPT. For each, we set the threshold so that at most 1% (and, separately, 0.5%) of verified human documents would be called AI, then confirmed that rate on 8,540 other human documents. At both thresholds, 0 of 223 held-out pipeline articles were called AI, while the same thresholds caught 13-80% of ordinary public AI text (Fakespot 80%, desklib 66%). The detectors even scored our articles as more human than real human writing, with AUROC between 0.22 and 0.40.

Two limits belong next to that result. A classifier trained on our pipeline's own output separates the same articles perfectly (AUROC 1.000), so a learnable fingerprint exists and general-purpose open detectors simply do not carry it. And commercial detectors such as GPTZero and Pangram were not part of the test, so the result says nothing about them.

Binoculars and Fast-DetectGPT are zero-shot methods: instead of learning from labelled examples, they score how predictable a text looks to a language model. Our 30 August audit found every zero-shot method we tried blind to humanized output, which overshoots the human unpredictability band (Binoculars about 1.06 for humanized output, against about 1.00-1.05 for human text).

Our own detector is measured on people as well. Across 9,963 verified human documents it called 4 AI. On a pool of 4,546 human documents spanning 17 kinds of writing, it called none (95% upper bound 0.08%). The cap we hold it to is 0.5% of human documents in every kind of writing, not just on average.

Public advice lands on material too. The Substack post 4 Ways to Make Your AI Content More Human (and More Searchable) warns that "Raw AI content often creates a 'sea of sameness,'" and argues that original data earns links from other websites. That is the author's opinion rather than a measurement, but it points where our section data points.

What the data supports is narrow. The open-detector pass reflects articles humanized to the level of a well-edited human writer. It is not evidence about first-hand records, and it does not make the articles untraceable. Separately, on one site, sections built on the company's own records read AI 37% of the time to our detector, against 71% for the rest, while a site that used general industry figures showed 98% against 100%.

We publish these limits, including the human writing our own detector gets wrong, because a pass rate without them means little.

Two notebooks side by side on a desk, one dense with handwritten proprietary data and the other sparse with generic text, illustrating the specificity gap between human-reading and AI-reading content
The difference between a section that reads human and one that doesn't often comes down to whether the records were collected before drafting began.

Why does AI-drafted content still read machine-like after humanizer passes?

Humanizing changes how text looks more than what it is built on. After our pipeline's humanizing, 144 of 146 September 2026 articles were still called AI by a detector trained on that pipeline's output.

The same articles did pass general-purpose detectors. At thresholds set on verified human text, at most 1 of the 146 was called AI, by the two zero-shot open detectors, and the two trained open classifiers flagged none. On 223 held-out articles, a classifier trained on our pipeline's output still separated them perfectly (AUROC 1.000), and commercial detectors such as GPTZero and Pangram were not tested. So "still reads machine-like" depends on the detector: the one that knows the pipeline recognises it, as of .

The humanizer market works on a simple theory: AI content reads robotic because of surface patterns, so replace those patterns. Swap flagged vocabulary, vary sentence length, add a contraction. A June 2026 Substack roundup of five humanizer tools judges them on whether they preserve meaning, improve flow, reduce robotic phrasing and save editing time, and its author cautions that "one good output does not prove consistency."

Call this the substance test: remove the brand name from a passage and ask whether it could appear word for word on a competitor's site. If it could, surface editing has little to work with. That is editorial guidance, not a measurement, and it does not predict a detector score.

The strongest evidence behind it comes from one site. For a sports equipment manufacturer whose AI-drafted articles drew on its own sales and shipping records, such as orders per state and models shipped per county, record-heavy sections read AI 37% of the time. The other sections of the same articles read AI 71% of the time. Same pipeline, same articles, different material underneath.

The limit is just as clear. For a payments company whose articles used general industry figures, the same comparison came out at 98% against 100%, a gap of two percentage points rather than thirty-four. We did not split all 2,625 sections by type of evidence; this is one strong site result with a clear counterexample.

Across body sections from all 146 articles, the sections that read human carried 7.9 numbers per 100 words, against 2.5 numbers per 100 words in sections that read AI. Our research does not read that as "use more numbers": the payments company's figure-rich articles barely differed. A plausible explanation is that figures only one company could publish did the work, not figures as such.

The problem practitioners describe may be two problems. One is packaging: the surface habits a model brings to any draft. The other is substance: text that anyone could have written, because it rests on information available to anyone. Humanizer tools are built for the first. Our data suggests, without proving it, that the second needs different inputs.

This article reports what we measured and where the limits are. The section scores cover 146 articles for 13 websites; the first-hand records comparison comes from one site, with the payments company as its limit. Where we offer practical guidance, we say so.

The next 12-24 months, scored

Where AI writing authenticity is headed

Three scored forecasts on how human-reading text, detectors, and humanizer tools evolve across the next two years.

27 sources analyzed7 community discussions4 industry publications3 newsletters2 blog posts
A

Three forecasts for AI writing and detection

Use these to weigh where original data beats stylistic polish before you pick a drafting tool or partner.

75/100
Medium confidence 12-24 months

As Google's Genesis, pitched to The New York Times, The Washington Post, and News Corp, and low-cost AI-interview workflows push drafting into mainstream newsrooms and creator workflows, the volume of homogeneous machine prose rises and the premium on differentiated first-hand writing grows.

Contrarian signal
63/100
Medium confidence 12-24 months

The common bet that humanizer tools let machine text reliably pass will falter over 12-24 months: because a detector trained on the output of a specific pipeline can still recognise that output after humanizing, surface rewriting from tools such as StealthGPT and GPTHumanizer plateaus while substance-grounded writing pulls ahead.

Weak signals watched: On one site, passages grounded in the firm's own sales records already read machine-made far less often than the writing around them, while general industry figures elsewhere showed a gap of only two points. Google pitched its Genesis news-writing tool to major publishers in 2023, and AI-interview products such as ContentPod start at $9.99 per month. Five-tool humanizer roundups and word-list 'tricks' now circulating focus on surface edits, while a detector trained on a specific pipeline's output still recognises that pipeline's humanized articles.

B

The data behind these calls

The public sources behind each forecast are shown so you can judge it for yourself.

Platform-embedded drafting floods the market with generic prose 75
Supporting evidence
  • Google wants you to let its AI bot help you write news articles is what puts this forecast on the board. [Industry Publication]Google is testing an AI news-writing tool internally titled Genesis, pitched to news organizations including The New York Times, The Washington Post, and News Corp (owner of The Wall Street Journal), according to three people familiar with… “Let it be said, journalists don't need Google to write their articles as 'a personal assistant.' And anything that Google (or any AI) could write has no real…”
  • The case rests on AI Interview Tool | Turn Voice Into Content - ContentPod. [Industry Publication]AI Interviews cost 5-20 credits by length: ≤5 min = 5 credits; 5-15 min = 10 credits; 15-30 min = 20 credits (ContentPod FAQ). “Voice-first requirement is a friction/limitation (cannot type answers), which the FAQ concedes.”
  • Getting interviewed by AI to write your technical content or blog posts is what puts this forecast on the board. [Industry Publication]The workflow originated from Daniela Petruzalek's Speedgrapher MCP demo for Gemini CLI, seen during a Google GDE call "a couple of weeks ago.". “the AI interviewer doesn't let you skip ahead. It asks open-ended questions that make you think outside your usual mental framework.”
Cosmetic humanizer tools hit diminishing returns 63
Supporting evidence
  • The case rests on How to Make AI Writing Sound Genuinely Human And Beat Top AI Detectors in 2026. [Blog]The author's team ran a content pipeline using Claude with system prompts covering E-E-A-T principles, citation standards, and formatting rules; output was strong for SEO/GEO but scored poorly on humanness. “LLMs are bad at 'find and replace' operations on their own output. They want to regenerate, not patch.”
  • Top AI Humanizer Tools for Content Creators Who Need Fast Draft is the strongest public backing for this call. [Substack / Newsletter]GPTHumanizer AI is ranked #1 / "Quick Verdict" top pick, described as free, no-sign-up, with a "Lite mode" for lighter cleanup and "deeper modes" for stronger rewriting. “This is about which tool fits the actual moment you are in when a draft needs to become publishable faster.”
  • Backing it: Has anyone solved the problem of making AI sound less "AI ish? [Community / Forum]jamboman_ tests daily using local LLMs and shell scripts on a Mac, and states plainly that evading AI detection is "extremely hard and much harder than people think.". “Lol people writing entire novels of instructions. You knowing you do that, all other instructions will be watered down? I just add 'in simple English' to my…”
C

What could flip these forecasts

Scenarios such as stronger base models or misfiring detectors that would restore the edge to surface-level humanizing.

A note on uncertainty

Treat these scores as weights, not verdicts. The top signal (95/100) leads on evidence, and the minority view (63/100) marks where sources spread out.

  • First-hand data becomes the authenticity marker. That is the first forecast to break if the regulatory or buying picture flips.
  • Cosmetic humanizer tools hit diminishing returns. Mounting evidence on the other side would move that one to the front.
Methodology Scores run 0-100 and weigh each signal by source authority, recency, and how many sources agree.

Have first-hand records but no time to write them up?

The AEO Content Engine turns your company's own records into articles written to read like expert human writing and structured so AI answer engines can cite them.

What should the evidence change about how you approach AI-drafted content?

Start with what the draft is built on. On one site, record-based sections read AI 37% of the time against 71% for the rest; general industry figures elsewhere showed 98% against 100%.

Humanizing and voice are separate jobs. As a Substack guide on humanizing AI writing puts it, "A humanizer strips out the machine tics. It doesn't give you a voice. Those are two different jobs." A pipeline that treats them as one job will serve one and neglect the other.

A practical checklist (guidance, not findings):

  • Collect first-hand records before drafting: your own counts, dates, places and outcomes, the facts only your company could publish.
  • Draft on that material, then edit the surface. Do not expect a rewrite pass to add specificity that was never collected.
  • Apply the substance test: remove the brand name and ask whether a passage could sit on a competitor's site.
  • Give extra attention to openings and fixed shapes such as decision frameworks, outlook charts and myth-versus-fact boxes, which read the most machine-like in our data.
  • Build author bios from a real profile rather than a generated summary.
  • Check a detector's false-positive rate on human writing before trusting its verdict, and do not treat a pass on one detector as a pass on all of them.

That sequence is where the AEO Content Engine starts: articles grounded in each company's own records, written to read like expert human writing and structured so AI answer engines can cite them. To see which passages of a draft read machine-written, run it through the free AEO Content AI Detector.

This article is part of our research series on how AI writes and how humans write. The overview of the whole series is How AI Writes vs How Humans Write.

Written by

Alex Shortov

CTO, AEO Content

Full-stack engineer and content infrastructure architect with 20 years of building enterprise systems.

Connect on LinkedIn

Summarize This Article With AI

Open this article in your preferred AI engine for an instant summary.

Frequently Asked Questions

What do AI detectors actually measure when they flag content?

Detectors differ, and most do not publish what they respond to. Trained classifiers learn from labelled examples, while zero-shot methods such as Binoculars and Fast-DetectGPT score how predictable text looks to a language model.

Our own detector combines a ModernBERT classifier with a separate predictability check from a pair of open Qwen2.5 models. It scores passages of about 300 words, calls a passage AI at 0.998 or higher and human below 0.5, and leaves the rest uncalled.

Is making AI content sound human mainly a writing or editing problem?

Our data points at sourcing as much as editing. ContentPod, an AI interview tool, argues on its product page that content creation "is usually an extraction problem," with the ideas stuck in calls, voice notes and half-finished docs. That is a vendor's framing, but it fits what we measured on one site, where sections built on the company's own records read AI far less often than the rest.

What is the difference between a humanizer tool and an author voice?

A humanizer rewrites surface habits such as stock transitions, puffery and repeated phrasing. An author voice is what the writing says and why only that person could say it: specific knowledge, first-hand observation and a consistent point of view. A humanizer can remove tics; it cannot supply the knowledge.

Does adding statistics to AI content make it read more human?

Not on their own. Body sections that read human carried 7.9 numbers per 100 words against 2.5 in sections that read AI, but on a payments company's site whose articles used general industry figures, the section comparison came out at 98% against 100%. The large gap, 37% against 71%, appeared on one site where the material came from the company's own records.

Does the detector misread human writing?

Rarely, and we publish the cases. Across 9,963 verified human documents it called 4 AI, and on a 4,546-document pool spanning 17 kinds of writing it called none (95% upper bound 0.08%). The human writing most often mistaken for AI in our tests was neutral institutional prose, such as government updates and policy summaries.

Why do article openings read more machine-like than later sections?

In our September 2026 data, 87% of first passages read AI, against 64-75% for later passages. Our research groups the double introduction, an opening paragraph followed by an intro teaser, with the other fixed rhetorical shapes that read the most machine-like. Why that happens is not something the data proves.

Can I test my own AI content for detection before publishing?

Yes. The free AEO Content AI Detector scores passages of about 300 words, shows a reading for each passage it scores, and reports origin and style as two separate readings. Checking a draft before and after adding first-hand material is a reasonable way to see what changed, though no single detector speaks for all of them.

Read next

Researcher analyzing AI detection calibration data and false-positive rates across writing registers on dual monitors

AI Detector False Positives: What a 0.5% Cap Really Means

Overhead view of an editorial desk with manuscript pages, statistical charts on a laptop screen, and handwritten analysis notes

How AI Detectors Work and Where They Break

Researcher reviewing printed documents and notebooks alongside a laptop showing an archive timestamp interface

How to Prove a Text Was Written by a Human