Back to blog
AI Detection

Does Turnitin Detect ChatGPT and AI Writing?

May 12, 20266 min read

Turnitin rolled out AI writing detection in 2023, and it's now one of the most widely deployed detectors in higher education. If you're a student, educator, or just curious how it works, here's the short version: Turnitin doesn't 'catch' AI text with certainty — it estimates a probability, and that estimate is noisier than most people assume.

How Turnitin's AI detector actually works

Turnitin's model looks at patterns in sentence structure, word predictability, and burstiness — the natural variation in sentence length and complexity that shows up in human writing. AI-generated text tends to be more uniform: sentences cluster around similar lengths, word choices skew toward the statistically 'safest' option, and paragraph rhythm can feel flat. The detector scores a document based on how much of it resembles that pattern.

Turnitin reports this as a percentage — for example, '38% of this submission is likely AI-generated.' That number is a probability estimate across the whole document, not a sentence-by-sentence fact. A single paragraph can be flagged heavily while the surrounding text isn't, especially in longer, mixed-authorship documents.

False positives are a real, documented problem

Multiple independent studies — including ones run by Turnitin itself — have found non-trivial false-positive rates, especially for non-native English writers, whose sentence patterns can be more uniform for reasons that have nothing to do with AI. Highly structured, formulaic writing (think: five-paragraph essays, lab reports, technical summaries) also tends to score higher on 'AI-likeness' even when a human wrote every word.

  • Uniform sentence length and rhythm reads as 'AI-like,' even from human writers
  • Formulaic structures (reports, standardized essay formats) score higher regardless of authorship
  • Non-native English patterns are statistically more likely to be flagged
  • Editing AI-assisted drafts by hand lowers — but doesn't reliably eliminate — the score

What actually changes a document's score

The most reliable way to lower an AI-detection score isn't a trick — it's variation. Detectors are built to notice uniformity, so text with a natural mix of sentence lengths, some structural asymmetry, and less predictable word choices tends to score as more human, regardless of how it was originally drafted. That's the same principle Humlexic's model is trained around: reworking AI-generated drafts so the rhythm and word choice look like normal human variation instead of a flat, average pattern.

If you're using AI tools as a starting point for a draft, running the result through a humanizer tuned for the detector you're being graded against — and then reading it over yourself — is a more dependable path than hoping the raw AI output passes.

Try Humlexic on your own draft

Paste your AI-assisted text and see it rewritten with natural rhythm and word choice — free to try.