AI

AI detectors in 2025: claimed accuracy versus real test results

Sep 3, 2025 · 12 min

Vendors claim up to 99.98% accuracy for AI detectors. Independent tests land at 79-86%, and over 61% of essays by non-native English speakers get falsely flagged as AI.

The market for AI detection is projected to grow from 0.58 billion dollars in 2025 to 2.06 billion in 2030, or 28.8% a year. Behind those numbers is a technology that promises more than it delivers.

The accuracy claims don't hold up

Winston AI claims 99.98% accuracy. In independent tests it reaches 79% precision with 6% false positives. Turnitin, the most widely used tool in education, claims 98% and delivers 80-86% in real-world tests.

The gap matters, because these tools are used to accuse people of cheating.

How the detectors work

AI detectors rely on four main methods.

Perplexity measurement looks at how predictable the text is. AI-generated text is typically more predictable and uniform, while human writing varies more.

Burstiness analysis measures variation in sentence length and structure. Humans mix short and long sentences; AI output is more even.

Linguistic pattern analysis looks at word choice and syntax, where AI leans on conventional patterns more than people do.

Embeddings and classification convert text into numerical representations that algorithms scan for AI-typical patterns.

The problem: simple paraphrasing can push detection accuracy from 99% down to under 10%.

The bias against non-native speakers

The most serious finding is systematic bias. Over 61% of essays written by non-native English speakers get falsely classified as AI-generated.

Run the numbers on a university with 75,000 assignments a year: a false-positive rate of 2-5% means up to 3,750 students falsely accused of cheating.

The biggest players

Originality.AI combines AI detection with plagiarism scanning in over 30 languages, claims 97-99% accuracy, and specializes in catching paraphrased AI content.

GPTZero was built by a Princeton student for education. Free up to 10,000 characters.

Copyleaks claims 99.1% accuracy with only 0.2% false positives.

Turnitin is the education standard, with the accuracy problems described above.

Danish universities are choosing another route

Danish universities are moving away from pure detection. From fall 2025, Copenhagen Business School allows generative AI in most take-home exams. The University of Copenhagen requires transparency about AI use instead of banning it.

Where detectors are actually useful

They have legitimate uses despite the limits:

  • Quality control of outsourced content in marketing and SEO
  • Initial screening, as a support tool rather than a final judge
  • Teaching the difference between AI-written and human-written text
  • Non-critical tasks like content organization and internal quality checks

Practical recommendations for 2025

  • Never use a detector as the only evidence
  • Be open about the limitations
  • Consider alternatives to detection
  • Account for the bias against non-native speakers
  • Focus on learning rather than punishment

Where this lands

AI detectors are neither miracle tools nor useless. But the numbers rule them out as judge and jury, and the Danish universities point at the workable path: transparency about AI use and human judgment, instead of chasing perfect detection.

Next step

Ideas are cheap.
Systems ship.

Tell me what you are building. I will tell you straight what is worth doing.

Start a conversation