Do AI Detectors Actually Work? What You Need to Know

Alex Turner·9 min read
do ai detectors work

You wrote every word of your essay yourself. Then a detector your professor ran it through comes back saying part of it is AI.

Can that really happen to an honest student? Yes. And it happens often enough that dozens of universities have quietly switched these tools off.

Here's the part almost nobody tells you first. Turnitin, the detector most schools use, says in its own official guidance that the score "should not be used as the sole basis for adverse actions against a student."

The company that sells the tool is telling teachers not to trust its number on its own. So do AI detectors actually work? Not reliably.

Here's what they really measure, how often they get it wrong, and what to do if one of them points at you.

What an AI Detector Measures

An AI detector is a tool that scans writing and guesses whether a person or a chatbot wrote it. That word, guesses, is the whole story. It doesn't read your essay for meaning.

It can't tell if your argument is smart or your facts are right. It just hunts for patterns.

There are two patterns it leans on:

  • Perplexity. How predictable your words are. If a computer can easily guess your next word, that reads as "AI," because a chatbot writes the words a model expects. Humans are supposed to be less predictable.

  • Burstiness. How much your sentences vary. People write in bursts: a long sentence, then a short one. A dense paragraph, then a punchy line. AI text comes out smooth and even, and that evenness reads as "machine."

Now the catch. Clean, careful, formal writing scores low on both.

If you were taught to write in plain, tidy sentences, if English is your second language and you stick to words you're sure of, or if you just edited your draft until it was neat, you look exactly like the thing the detector is trained to catch.

It can't tell a careful human from a machine. It was never built to.

How Accurate Are AI Detectors

Start with the most telling example. OpenAI, the company that makes ChatGPT, built its own AI detector in 2023, then shut it down six months later for being too inaccurate.

If the makers of ChatGPT can't reliably spot ChatGPT, that tells you where this technology really is.

And as of 2026 it still hasn't shipped a replacement, even though it reportedly built a text-watermarking tool and chose not to release it.

Advertisement

Then there's the test that went viral. People ran historical documents through these tools, and one rated the U.S. Constitution as mostly AI.

It was written by hand in the 1700s. GPTZero's founder, Edward Tian, explained why: old, famous documents show up constantly in the text chatbots learn from, so the AI writes in their style, and the detector flags that style as AI.

Plenty of good student writing is formal in the exact same way. The most serious study came out of Stanford.

Researchers fed 91 essays by non-native English speakers into seven popular detectors, and the tools flagged most of them as AI. Every single essay was written by a human, and about one in five got flagged by all seven at once.

Here's the record, side by side:

Tool or study

What happened

OpenAI's own detector

Shut down in 2023. Caught just 26% of AI text and wrongly flagged 9% of human text.

Stanford, 7 detectors

Flagged 61% of essays by non-native English speakers. All were written by humans.

Turnitin

Admits scores can be off by up to 15 points, and can't be the sole basis for punishment.

ZeroGPT

Rated the U.S. Constitution ~92% AI and the Declaration of Independence higher.

Grammarly (2025 test)

Editing a human essay with Grammarly pushed its Turnitin AI score to 16%.

Turnitin doesn't even claim otherwise. Its own guidance is blunt about how much weight its score deserves:

Turnitin's AI writing detection "should not be used as the sole basis for adverse actions against a student."

And independent research is harsher. One academic study found false alarms on real human writing running from 15 to 26 percent. That's up to one honest paper in four.

You'll see detectors advertise 98 or 99 percent accuracy. Those are the companies' own numbers.

When independent researchers test them, the picture falls apart. A March 2026 benchmark of eight detectors found not one topped 80 percent on real writing. Edit that text even a little and accuracy can drop as low as 20 percent.

The marketing number and the real-world number are not the same thing.

Why Universities Are Turning AI Detectors Off

You don't have to take a blog's word for any of this. Look at what universities did once they tested the tools for themselves.

Turning them off is not universal. Many colleges still run AI detection (one 2026 estimate puts it near 40 percent of four-year schools), which is exactly why this can still land on you. But the schools that studied the tools closely keep reaching the same verdict.

Vanderbilt disabled Turnitin's AI detector back in 2023, and the math was simple. It had sent about 75,000 papers through Turnitin in a single year, so even a 1% false-positive rate meant roughly 750 students wrongly flagged. It decided the tool wasn't worth that.

Advertisement

School

Turned it off

Why

Vanderbilt

2023

Too many likely false positives; no insight into how it works.

University of Pittsburgh

2026

Use "is simply not supported by the data."

University of Waterloo

2025

Found cases where it called human writing 100% AI.

Yale, Johns Hopkins, Northwestern

By 2026

Accuracy and fairness concerns; among 50+ schools that switched it off.

Pittsburgh didn't soften it when it dropped the feature:

[The tool] is simply not supported by the data and does not represent a teaching practice that the Teaching Center can endorse or support.

At Harvard, the professors who deal with suspected AI in person are just as doubtful. Jeffrey Miron, who runs undergraduate economics, said the tools aren't the answer:

If it's zero percent, I don't know if I believe it. If it's 100 percent, I might have believed it, but I don't think those detection programs are a cure-all.

His colleague Mary Lewis, who runs undergraduate history, named the practical problem: even when a teacher suspects AI, it's

virtually impossible to prove, and so I think faculty are hesitant... because they think it will be a waste of their time.

Their fix hasn't been a better detector. It's been to move more writing into the classroom, where they can watch it happen.

Who Gets Wrongly Flagged

This is the part that should bother you most. The tools don't just miss. They miss in one direction, and it's the unfair one.

The students most likely to be wrongly flagged are the honest ones:

  • Non-native English speakers, for the reason the Stanford study laid out: simpler words and steadier sentences look like a machine to these tools.

  • Students who write formally or use Grammarly. That 2025 test found Grammarly's edits alone pushed a real human essay to a 16% AI score.

  • Students on the autism spectrum, whose writing can be more literal and structured, which reads as low variety to a detector.

On that last point, the cases are real. A student at Adelphi University who is autistic sued after a professor cited a "100% AI" report on an essay he wrote himself; other tools rated the same essay human.

A student at the University of Michigan filed a disability discrimination claim over a similar accusation.

Meanwhile, the students who really cheat often walk right through. There's a whole class of tools called "humanizers" that rewrite AI text to beat detectors.

So the real cheater pastes their essay into a humanizer and passes, while the honest student who wrote plainly gets called in. The tool punishes the wrong person.

Advertisement

What to Do If You're Falsely Accused

Because these tools are wrong often enough, treat a false flag as something that could happen to you. The good news: you can build your own proof as you go.

  • Write where your history is saved. Google Docs and Microsoft Word both keep version history. That slow record of you typing, deleting, and rewriting over days is the strongest evidence a human wrote it. A detector score is a guess. Your edit history is a receipt.

  • Keep your drafts and notes. Don't trash the outline or the messy first version. If anyone asks, that's what clears you.

  • Don't panic, ask about the process. A score isn't a verdict. Point to Turnitin's own line that it can't be the only basis for punishment, and ask what your school's academic integrity policy requires.

  • Get a second opinion. Run the same text through a couple of other detectors. They usually disagree with each other, and that disagreement is exactly the point.

Wondering whether this follows you into the admissions process? That's a separate question, and we answer it in Do College Admissions Check For AI?.

The Bottom Line

AI detectors don't work the way their marketing says, and the schools and companies closest to them keep admitting it.

They measure how predictable your writing is, not whether you're honest, and careful writers are predictable all the time.

None of that is permission to let a chatbot do your work. If your school bans it, using it is a real risk no matter what any detector says, and the point of the assignment is the thinking you'd be skipping.

Use AI like a tutor, to help you understand, not like a ghostwriter. Then write where your history is saved, keep your drafts, and you'll have the one thing a guessing machine never will: proof.

Frequently Asked Questions

Can Turnitin detect ChatGPT? It tries, and sometimes it's right. But it works by spotting patterns, not by knowing, so it also flags plenty of human writing. Turnitin itself says the score shouldn't be the only reason a student is punished.

Can AI detectors be wrong? Often. OpenAI's own detector caught just 26% of AI text before it was shut down, and a Stanford study found detectors wrongly flagged 61% of essays by non-native English speakers. Real human writing gets flagged all the time.

Does using Grammarly trigger AI detectors? It can. A 2025 test found that editing a human essay with Grammarly pushed its Turnitin AI score to 16%. Basic spelling and grammar fixes are usually fine, but heavy rewriting suggestions can nudge the number up.

How do I prove I didn't use AI? Write in Google Docs or Word with version history on. The timeline of you building the document over hours or days is the clearest proof a human wrote it. Keep your outline, drafts, and sources too.

Which colleges have turned off AI detection? Vanderbilt, MIT, Yale, UCLA, Georgetown, the University of Pittsburgh, and the University of Waterloo are among the 50-plus schools that have disabled Turnitin's AI detector, though many colleges still use it.

Advertisement

Related Articles