ContentScanContentScan
Log inSign up free
RESEARCH

The State of AI Content Detection in 2026: What the Research Actually Shows

Ahmad Raza··8 min read

Most “AI detector accuracy” claims you’ll see on comparison pages — including some of our own elsewhere on this site — are marketing numbers with no methodology attached. This page is different: it’s a roundup of what independently published research actually found, with a link to the primary source for every figure below. Where a number comes from a vendor’s own marketing rather than independent testing, we say so explicitly.

How much of the web is actually AI-written?

Graphite.io ran a large-scale study sampling 43,000 URLs from the CommonCrawl archive, filtering for English-language articles published between January 2020 and May 2025. Using a 500-word-chunk AI classifier (validated against a 4.2% false-positive rate on pre-ChatGPT articles and a 0.6% false-negative rate on known GPT-4o output), they tracked what share of newly published articles were majority AI-generated over time.

  • November 2022 (ChatGPT launch): near-zero baseline.
  • November 2023: roughly 39% of newly published articles were majority AI-generated.
  • From mid-2024 onward: growth plateaued, holding relatively stable rather than continuing to climb.

The plateau matters as much as the growth curve — it suggests the “AI will write 90% of the internet” framing you’ll see in some headlines has outpaced what the data actually shows. Read the full methodology at graphite.io.

False positive rates: independent testing vs. vendor claims

This is where detector marketing and independent research diverge the most. GradPilot’s 2026 comparison distinguishes between an independent 2025 study run by the University of Chicago Booth School of Business and the false-positive rates detector vendors publish about their own products — a distinction most comparison pages don’t bother to make.

DetectorIndependent test resultVendor self-reported claim
PangramEssentially zero on long/medium text; ≤1% on short passages~0.01% overall (self-reported)
GPTZero≤1% on medium/long text; up to 2.4% on short passages
Originality.ai≤1% on medium/long text; up to 2–3% on short passages
Turnitin~4% at sentence level (third-party analysis)<1% at document level (vendor claim)
CopyleaksNot independently benchmarked in this study~0.2% (1 in 500), vendor claim

The gap between Turnitin’s document-level claim and the sentence-level third-party figure is a useful illustration of why we don’t publish a single self-rated accuracy percentage for ContentScan on this site — see our detector comparison guide for how we approach that instead. Full detail at gradpilot.com.

The ESL bias problem is real and well-documented

The most consequential finding in this space is a 2023 peer-reviewed study by Liang, Yuksekgonul, Mao, Wu, and Zou, published in the Cell Press journal Patterns. Testing seven widely used AI detectors against 91 TOEFL essays written by non-native English speakers, the researchers found that more than half of the non-native-authored essays were incorrectly flagged as AI-generated — while the same detectors scored near-perfect accuracy on essays written by US eighth-graders.

The researchers’ hypothesis: non-native English writing tends to have lower “perplexity” — less lexical variability — which is exactly the statistical signal most detectors use to flag AI-generated text. The bias isn’t a bug in one product; it’s a structural weakness in how perplexity-based detection works, which is why sentence-level review matters more than a single score for any writing sample from a non-native speaker.

Read the full study (open access) on arXiv, or the published version via Patterns / ScienceDirect.

What this means if you’re choosing a detector

  • Treat every accuracy claim as unverified until you see the methodology. “99%+ accuracy” with no linked study is a marketing number, not a research finding.
  • Independent, third-party testing and vendor self-reporting are not the same evidence. Ask which one you’re looking at.
  • If you’re evaluating ESL or non-native writing, weight false positives heavily. The bias documented by Liang et al. hasn’t been fully solved industry-wide, even in 2026.
  • Prefer sentence-level detail over a single score. It lets a human reviewer catch exactly the kind of false positive this research describes, instead of trusting one number blind.

That last point is the design decision behind ContentScan’s own AI content detector: a sentence-by-sentence confidence breakdown rather than a single opaque percentage, precisely because the research above shows why a single score can mislead.