AEO Content AI: free AI content detector AEO Content AI: free AI content detector

Which Parts of an AI-Written Article Give It Away

An AI-written article gives itself away through the shape of its sections more than through individual words.

Printed article sections marked with color-coded annotations showing different structural patterns on an editor
Three things content marketers believe about AI detection. Myth or fact?
Call each one, then see how other readers called it.
1 Scanning for overused AI words is the most reliable way to spot machine-written content.
2 The opening of an AI article reads more machine-like than the sections after it.
3 Most editors can tell AI-written content from human writing accurately just by reading it.

Quick Answer

An AI-written article gives itself away through the shape of its sections more than through individual words. In our scoring of 2,625 sections from 146 AI-drafted articles, decision frameworks read AI 100% of the time and opening paragraphs 86%, while author bios built from a real profile read AI only 8% of the time. Reused phrases made up 0.5% or less of every prose section, so vocabulary was not what separated them. The detector we used, a ModernBERT classifier with a Qwen2.5-based predictability check, was trained on that pipeline's output.

Did this answer your question?

What Gives an AI-Written Article Away?

The section shape, more than individual words. In our study, detection rates varied far more by section type than by vocabulary, even inside a single article.

A resource list in an AI-drafted piece read AI 42% of the time and read human 21% of the time. A call-to-action block read AI 36% of the time and read human 33% of the time. In the same articles, decision frameworks read AI 100% of the time and opening paragraphs 86%. An article is judged by its most machine-like passage, not by its average.

Our September 2026 analysis scored every section of 40 words or more separately, across 146 articles written for 13 websites. Reused phrases shared across websites made up 0.5% or less of every prose section, yet 81-100% of every generated prose section type still read AI. Earlier de-templating work had already removed the shared phrases. What remained were fixed rhetorical shapes: decision frameworks, outlook charts, before-and-after comparisons, myth-vs-fact boxes, key takeaways and the double introduction.

Of the 146 articles, 56 contained at least one passage that read human, but 144 were still called AI at the article level, because a single machine-like passage is enough for that call.

This article covers which sections read most machine-like, what our human controls showed, why first-hand data changes the result, and where the signal is heading over the next 12 to 24 months. Every figure comes from that internal study or from named external sources.

The giveaway in an AI article is less a word than a shape. The opening teaser, the myth-vs-fact block, the decision framework with its tidy options and recommendation: in our data these fixed rhetorical moves read machine-like far more consistently than any vocabulary did.

Writers who study this from the reading side describe the same thing. John R. Gallagher calls large language models genre machines, systems that reproduce the conventions of the genres they learned. Practitioners on r/content_marketing report that AI articles "humanized" by synonym replacement still read AI. The shape survives the word swap.

Why would the shape persist? One plausible explanation is instruction tuning, which rewards certain answer formats until a model reaches for them by default: an opening that sets broad context, parallel supporting sections, an outlook that ends on a confident claim. Our data cannot prove that mechanism. It shows where the shapes concentrate, and that removing shared phrasing did not remove the signal.

Want fewer template-shaped sections in your articles?

The AEO Content Engine starts from your company's own records instead of generic figures, so sections carry detail no template supplies, built for AI answer engines to cite.

Forward Signal - 12-24 months horizon

How AI-Written Text Will Give Itself Away Next

Three scored forecasts on what will separate machine-drafted from human writing as spotting methods mature over the next one to two years.

27 sources analyzed10 community discussions3 blog posts1 academic source1 newsletter
A

What will mark machine-drafted writing next

Use these to judge which signals of machine authorship will still hold and which will fade as writers and tools adapt.

68/100
Medium confidence 12-24 months

Over the next one to two years the durable marker of machine-drafted writing will be structural (repeating list scaffolds, formulaic openers such as 'in today's fast-paced world' used four times in 800 words, and uniform paragraph rhythm) rather than a fixed list of banned words, because word blacklists contradict each other and flag ordinary human writing.

Contrarian call
64/100
High confidence 12-24 months

Confident 'I can tell instantly' spotting will stay unreliable through the forecast window; with unaided accuracy at 50-52% and automated tools that practitioners call hit-or-miss, buyers and educators will lean less on gut judgment and more on structured, repeatable analysis.

Faint signals worth tracking: Guides claiming to name the giveaway words share almost no overlap, and the phrases they flag ('significantly', 'let's take a look at', 'ultimately') turn out to be common in genuine human writing. Study participants distinguished human from machine text at 50-52%, no better than chance, and equally poorly across dating, hospitality, and professional samples. Machine writing's recurring weakness is generalization without specifics and loss of coherence in longer pieces, and passages grounded in proprietary records read human far more often.

B

Evidence for and against these calls

Supporting studies and skeptical practitioner accounts are both listed so you can weigh each forecast yourself.

Lived detail as the divider 95
Supporting evidence
Structure over vocabulary 68
Supporting evidence
Intuition keeps missing 64
Supporting evidence
  • Was this written by a human or AI? | Stanford HAI is what puts this forecast on the board. [Academic]Study participants could distinguish between human- and AI-generated text with only 50-52% accuracy, roughly the same as a coin flip (Hancock et al.). “That's worrisome because it creates a risk that these machines can pose as more human than us.”
  • Backing it: How can I tell if content is written by AI tools? [Community / Forum]The original poster (OP) is a client working with a freelance writer, seeking to verify content is human-written and not produced by ChatGPT or other generators; the OP explicitly notes awareness that "AI-detection tools exist, but many… “AI will force it in awkwardly or ignore it completely. A real writer will either do it naturally or push back and ask why.”
C

What could flip these forecasts

Shifts in model behavior, writer habits, and detection accuracy that would change how machine writing reveals itself.

Before you rely on these numbers

No forecast here is a sure thing. The strongest signal scores 95/100; the minority read (64/100) exists because sources weigh the trend differently.

  • Lived detail as the divider. That call weakens first if regulators or buyers move in the opposite direction.
  • Intuition keeps missing. That one becomes the more durable forecast if the source mix shifts toward stronger contrary evidence.
Methodology Each signal scored 0-100 by an evidence-weighted model based on source authority, recency, and how many sources support it.

Does the Shape Signal Also Flag Well-Organized Human Writing?

Rarely. Of 4,546 verified human documents scored by our detector, none were called AI, and across 9,963 verified human documents in total it called 4 AI (0.04%).

That result matters because the obvious counterargument is straightforward: maybe the detector just penalizes organized prose. A skilled human essayist who writes clean, structured paragraphs with consistent transitions could, in theory, look like whatever the model has learned. If that were true, the section findings above would be useless.

We tested that worry directly. On the 4,546-document pool, spanning 17 kinds of writing, none scored 0.998 or above. Across the full set of 9,963 verified human documents, including pre-2021 Medium posts and US federal agency prose, it called 4 AI. The human writing it did misread was formal and institutional: the previous version of the detector called 8 of 637 agency documents AI before we retrained it with 100 public-domain agency documents, and the current version calls 2 of the same 637.

We also ran control tests using human-authored chunks from Medium posts. Passages of 100, 200 and 300 words were called AI 0% of the time, and 50-word chunks 1.0%. Short human passages were not flagged just for being short, which is what lets us score sections one at a time.

The contrast with unaided human judgment is sharp. According to Hancock et al. (Stanford HAI, March 2023), people asked to distinguish human from AI text in dating, professional and hospitality profiles did so with 50-52% accuracy. That is not meaningfully different from a coin flip.

The same study shows why. Participants leaned on cues such as grammatical correctness, first-person pronouns, references to family life and informal language, and wrongly read them as human. Those cues are easy for a model to produce and common in human writing, so they cannot separate the two.

The practical consequence runs in both directions. An editor who reads an AI-drafted section and feels it "seems fine" is working from a signal that did not beat chance in controlled settings. An editor who reads a tidy human section and suspects AI has no better basis. Confident eyeball judgment is close to a coin flip. A measured reading, with its false-positive rate stated, is a better basis for the call.

Why Do Sections Built on First-Hand Data Read Human More Often?

On one manufacturer's site, sections built on its own sales and shipping records read AI 37% of the time, against 71% for the other sections of the same articles.

This came out of the same September 2026 scoring, looking inside one website's articles. That company's articles drew on its own records, such as orders per state and models shipped per county, and the record-heavy sections read AI about half as often as the rest. It is the most actionable finding from the study.

The effect has a limit. For a payments company whose articles used general industry figures, the same comparison came out at 98% against 100%, essentially no difference. Numbers alone did not do it. First-hand records did.

Across all body sections, the ones that read human carried 7.9 numbers per 100 words, against 2.5 in sections that read AI. That is a correlation, not a recipe: the payments comparison shows that adding general figures does not move the reading. The numbers that mattered were ones only that company could publish.

A plausible reason is that a paragraph built around a specific count from a company's own database is harder to fit into a default template. Our data shows the correlation. It does not prove that mechanism.

Reused phrases shared across websites made up 0.5% or less of every prose section. Phrase-level copying was essentially gone, yet generated prose still read AI, which is why we look at section shapes and evidence rather than at word choice.

I find this distinction practically important because most editing advice targets the phrase level: replace overused words, vary sentence openers, remove clichés. In our data that level was already clean, and all 16 decision frameworks still read AI. The lever that moved sections was evidence the model could not have written from general knowledge.

A review order for AI drafts, based on where the tells sat

The section scores suggest an order of work for anyone editing AI-assisted drafts. What follows is guidance drawn from where the machine-like passages concentrated in our 2,625 sections, not a separately measured procedure.

  1. Start with the opening. First passages read AI 87% of the time, against 64 to 75% for later passages. Replace a context-setting opener with the most specific fact the article has, ideally one only your organization could publish.
  2. Question every fixed-format block. Decision frameworks read AI in all 16 cases, outlook charts 99% of the time, before-and-after comparisons 94% and myth-vs-fact boxes 87%. Keep a block only if it carries information readers need, and build it from your own cases rather than a generic template.
  3. Check for a double introduction. Intro teasers read AI 87% of the time and opening paragraphs 86%. Two stacked introductions give the same template two chances to show.
  4. Bring in first-hand records before rewriting sentences. On the one site built on its own sales and shipping records, those sections read AI 37% of the time against 71% for the rest. General industry figures did not move the reading (98% against 100%).
  5. Use real people and real choices. Author bios built from a real profile read human 51% of the time, calls to action 33% and curated resource lists 21%. A named author with a real background and sources a person actually chose are parts of the page the model did not write from scratch.
  6. Leave word lists for last. Shared phrases were already down to 0.5% or less of each section, and the prose still read AI. Word-level cleanup is worth doing, but it is not the step that changes the reading.

None of these steps makes a draft untraceable, and they are not meant to. A detector trained on a specific pipeline can still recognize that pipeline's output. The point is a better article: one built on evidence a reader cannot find anywhere else.

What Will Determine Whether AI Content Reads Human in the Next 12-24 Months?

We expect fixed section shapes to stay a stronger tell than vocabulary, first-hand evidence to stay the main lever, and unaided reading to stay unreliable. None of this is certain.

  • Prediction: Structural shape outlasts vocabulary as a marker.

    Our data already points that way. Shared phrases were down to 0.5% or less of each section, yet every generated prose section type read AI 81-100% of the time. Word lists also misfire in the other direction: one writer on Medium found that many phrases popular guides flag as AI are ones real writers use all the time.

    Weak signal: Guides that promise to name the giveaway words end up flagging ordinary human phrasing. Why it matters: Teams still editing for word choice rather than section shape will keep producing drafts that read machine-written.

  • Prediction: Specific, verifiable detail becomes the clearest dividing line.

    On one manufacturer's site, sections built on its own sales and shipping records read AI 37% of the time, against 71% for the rest of the same articles. Where articles used general industry figures, the split was 98% against 100%. Detail that only one company could publish is harder to produce from general knowledge than polish is.

    Weak signal: Models can generate plausible-sounding specifics, but invented figures fail as soon as someone checks them. Why it matters: Evidence sourcing becomes a content quality decision, not just a research step.

  • Prediction: Intuition-based review stays unreliable through the forecast window.

    In the Stanford study, readers told human from AI text with 50-52% accuracy and wrongly read grammatical correctness, first-person pronouns and informal language as signs of a human writer. Those cues are easy for a model to produce and common in human writing.

    Weak signal: Rising public awareness of AI writing does not by itself give readers a better cue to rely on. Why it matters: Editorial gut calls on authorship sit close to chance in controlled tests, however experienced the reviewer feels.

One open question is whether newer models will shed these default shapes. Nothing in our September 2026 data settles that. It shows the shapes are there today, and that removing shared phrasing did not remove their signal.

Two document pages side by side showing contrasting structural patterns: uniform and formulaic versus varied and annotated
The same detection score can appear twice in the same article. Structural consistency, not vocabulary, is what separates the sections that flag from the sections that do not.

How We Measured Which Sections Flag as AI-Written

We scored 2,625 visible sections across 146 AI-drafted articles and found a 92-point spread between section types inside the same articles.

Between September 1 and 13, 2026, our pipeline wrote 146 long-form articles for 13 different websites. Every visible section of 40 words or more was scored separately with our AI detector, a ModernBERT classifier paired with a separate predictability check running on a pair of open Qwen2.5 models. Passages scoring 0.998 or above were called AI, passages below 0.5 were called human, and everything in between was not called. The detector was trained on our pipeline's earlier output, so it recognizes that pipeline's fingerprint by construction. This is an internal evaluation of our own articles, not a general detector benchmark, as of .

Across the articles, 73% of passages read AI, 16% were not called, and 11% read human. That last number is not nothing: more than 1 in 10 passages in our AI-drafted articles scored below the human threshold. The question is which ones.

Section typeSectionsRead AIRead human
Decision framework16100%0%
Prediction or outlook chart9499%0%
Before and after comparison9994%2%
What will matter most (outlook)10289%5%
Intro teaser14487%3%
Closing paragraph14287%5%
Myth vs fact13587%3%
Key takeaways10287%2%
Opening paragraph14686%3%
FAQ14481%4%
Body section76081%7%
External resources list5242%21%
Call-to-action banner12536%33%
Author bio398%51%

The answer is sharp. First passages read AI 87% of the time, compared with 64 to 75% for later passages in the same articles. Decision frameworks read AI in all 16 cases, outlook charts 99% of the time and before-and-after comparisons 94%. At the other end, author bios built from a real profile read AI only 8% of the time and read human in 51% of cases.

The position effect deserves a name: call it the opening penalty. The first passage of an article is where a default frame is most likely to sit, a context-setting claim about why the topic matters and a smooth transition into the body, and it is the part that read most machine-like in our data.

Practitioners describe the same pattern. One commenter on r/content_marketing put it plainly: "AI tends to be super formulaic - intro, 3 main points, conclusion, every single time. Real writers are messier."

Human-reading passages do not rescue an article. 56 of 146 articles contained at least one, yet 144 of 146 were still called AI overall, because the article-level call follows the most machine-like passage.

The practical consequence is specific: if you are editing an AI draft, the opening and the fixed-format blocks are where the tells concentrate. Fixing word choice elsewhere while leaving those untouched is the wrong order of operations.

The shape signal held across our September 2026 sample. On the one site whose articles were built on its own sales and shipping records, record-heavy sections read AI 37% of the time against 71% for the rest.

What changes when you know this: the checklist for reviewing an AI draft shifts from vocabulary to architecture. The question stops being "does this word sound AI?" and starts being "is this section a default template, and does it carry anything only we could publish?" Those are different questions and they lead to different edits.

For content teams using AI assistance at scale, that makes evidence sourcing the first step, before editing. That is where the AEO Content Engine starts: articles grounded in each company's own records, written to read like expert human writing and structured so AI answer engines can cite them. To see which passages of a draft read machine-written, run it through the free AEO Content AI Detector.

This article is part of our research series on how AI writes and how humans write. The overview of the whole series is How AI Writes vs How Humans Write.

Written by

Alex Shortov

CTO, AEO Content

Full-stack engineer and content infrastructure architect with 20 years of building enterprise systems.

Connect on LinkedIn

Summarize This Article With AI

Open this article in your preferred AI engine for an instant summary.

Frequently Asked Questions About Identifying AI-Written Articles

What does AI writing actually look like structurally?

Fixed formats repeated across sections. In our scoring of 146 AI-drafted articles, decision frameworks read AI in all 16 cases, outlook charts 99% of the time and before-and-after comparisons 94%, while reused phrases were 0.5% or less of each section. Writer John R. Gallagher describes large language models as genre machines that reproduce the conventions of what they learned.

Should I trust my own reading to judge whether a draft is AI-written?

Not as the primary check. In a Stanford study, people judging dating, professional and hospitality profiles told human from AI text with 50-52% accuracy, and they wrongly read grammatical correctness and first-person pronouns as signs of a human writer. A section can feel natural to read and still read machine-written to a measured check.

Why do some sections of an AI article read more human than others?

Two things stood out in our data. Sections the model did not write from scratch read human more often: author bios built from a real profile 51% of the time, calls to action 33% and curated resource lists 21%. And on one site, sections built on the company's own records read AI 37% of the time against 71% for the rest.

What is the most effective edit to make an AI section read more human?

Change the evidence before you change the words. On the one site where articles used the company's own records, those sections read AI 37% of the time against 71% for the rest, while general industry figures made almost no difference (98% against 100%). The edit that matters is upstream: data only that company could publish.

Can one human-written passage make an AI draft read human overall?

Not in our data. 56 of the 146 AI-drafted articles contained at least one passage that read human, yet 144 of the 146 were still called AI at the article level, because a single machine-like passage is enough for that call. Improving one section helps that section; the fixed-format blocks and the opening still need their own work.

Read next

Researcher analyzing AI detection calibration data and false-positive rates across writing registers on dual monitors

AI Detector False Positives: What a 0.5% Cap Really Means

Overhead view of an editorial desk with manuscript pages, statistical charts on a laptop screen, and handwritten analysis notes

How AI Detectors Work and Where They Break

Researcher reviewing printed documents and notebooks alongside a laptop showing an archive timestamp interface

How to Prove a Text Was Written by a Human