AI Content Detector Accuracy Is Worse Than You Think
I Tested 13 AI Content Detectors on Camera
The test used three writing samples on the same topic. One was an article I wrote myself in May 2020, years before ChatGPT existed.
The second was ChatGPT writing on the same topic with a plain prompt.
And the third was ChatGPT again. But this time it was trained on my own article first, and asked to match my tone and writing style.
I didn’t run the tests ahead of time or know what would happen.
As I said going in: “My hypothesis is that they don’t [work]. They’re easily fooled by anyone who knows how to prompt AI properly, and they also spit out false positives, which land freelancers in hot waters with clients.”
The Results, Tool by Tool
| Tool | My 2020 article | Raw ChatGPT | ChatGPT + my voice | Score |
|---|---|---|---|---|
| Winston AI | Correct | Correct | Correct | 3/3 |
| Originality.ai | Wrong (50-56% AI) | Correct | Correct | 2/3 |
| Surfer SEO | Correct | Correct | Wrong | 2/3 |
| IsGen AI | Correct | Correct | Wrong | 2/3 |
| GPTZero | Correct | Correct | Wrong | 2/3 |
| Copyleaks | Correct | Correct | Wrong (40% AI) | 2/3 |
| Scribbr | Correct | Correct | Wrong | 2/3 |
| QuillBot | Correct | Correct | Wrong | 2/3 |
| BrandWell (Content at Scale) | Correct | Wrong | Wrong | 1/3 |
| Undetectable AI | Wrong (99% AI) | Correct | Wrong | 1/3 |
| Grammarly | Correct | Wrong (41% AI) | Wrong | 1/3 |
| ZeroGPT | Correct | Wrong | Wrong | 1/3 |
| AI Detector.com | Correct | Wrong | Wrong | 1/3 |
AI content detector accuracy ranged from perfect to worse than a coin flip, depending entirely on which of these 13 tools you picked.
What Caught the AI Content and What Didn’t
Winston AI was the standout. It flagged my human writing as 99% human. The raw ChatGPT article scored 98% AI. And the voice-trained article came in at 76% AI.
That third result is the one that matters. 10 of the other 12 tools missed it completely.
I said it in the video: “I think Winston AI might be the best candidate for an AI content detection tool. But with an error margin of 24%, it’s still not something that I would rely on.”
The bottom of the table is where it gets uncomfortable.
BrandWell, still branded Content at Scale at the time, called all three samples human. That includes two that ChatGPT wrote start to finish.
AI Detector.com did the same thing. I called it out on camera: “That means Brandwell has a useless AI content detection tool.”
Undetectable AI was worse in a different direction. It markets itself with logos from Forbes, Buzzfeed, and Business Insider and claims 18.5 million users.
And it flagged my own human 2020 article as 99% AI-generated while missing the voice-trained AI article entirely.
That’s the false positive that costs a freelancer a client relationship over content they wrote with their own hands.
Why None of This Matters as Much as You Think
Even the best-performing detector shouldn’t be the reason you accept or reject a piece of writing.
Google’s own developer guidance on AI-generated content points to “high-quality, people-first content.” And it says that standard applies “however content is produced.”
The standard is quality.
Google has been building AI products since long before ChatGPT existed. And they ship Gemini and AI Overviews into search results by default now.
So the idea that they’d penalize AI-assisted writing at the same time doesn’t hold up.
Want this done for you? Book a free strategy call →
The only useful question is whether the finished piece is worth reading.
A cent-a-word content mill producing keyword-stuffed garbage writes worse copy than someone who knows how to direct AI. No detector percentage changes that.
A Year Later, I Retested BrandWell
Out of curiosity, I ran BrandWell’s free checker again this week. The samples were fresh, with nothing to do with the ones in the video.
I fed it four samples:
- A real human paragraph from a post already live on this blog
- A paragraph of raw AI writing stuffed with clichés
- A second AI paragraph with no obvious tells
- A human paragraph that uses the word “leverage” twice on purpose, just to see if vocabulary alone would trip it up
This time it caught both AI paragraphs correctly and passed both human paragraphs as human. Four for four.
That’s a real change from the tool that missed two out of two AI samples in last year’s test.
Whether that’s a genuinely improved model or just an easier set of samples, I can’t say for certain.
But a detector’s accuracy today doesn’t tell you much about six months from now, in either direction.
The tool that embarrassed itself on camera last year isn’t necessarily the same tool running under that name today.
What This Means for Your Marketing
If a client or job posting requires AI detection clearance, check which tool they’re using.
It could be BrandWell, which called ChatGPT drafts human, or Undetectable AI, which called my own writing a machine’s. Neither outcome tells you anything reliable about quality.
I’d check whether the claims in the copy are true. That matters more than any AI score.
If you want that kind of check, run your draft through the free fact-checker I built.
It won’t score you on AI-likeness. But it will catch the unsourced stats and shaky claims that get copy rejected for a better reason.
If it’s rhythm rather than facts you’re worried about, how to make AI writing sound human is the one to read. It covers the six patterns that give AI away, regardless of what any detector says.
FAQ
How accurate are AI content detectors?
In my test of 13 tools against three known samples, only one, Winston AI, correctly identified all three. Most tools caught obvious ChatGPT output but missed AI writing trained on a specific voice, and two of them flagged genuine human writing as AI-generated.
Which AI content detector is the most accurate?
Winston AI performed best in my test, correctly leaning AI on a sample written by ChatGPT but trained on my writing style, something 10 of the other 12 tools missed entirely.
Is Copyleaks accurate for detecting AI content?
Copyleaks caught my human writing and a raw ChatGPT article correctly, but only flagged 40% of a voice-trained AI article as AI-generated, effectively calling it majority human when it wasn't.
Does Content at Scale's AI detector still work?
The company rebranded to BrandWell. In my original test it called all three samples human, including two written entirely by ChatGPT. When I retested it a year later on new samples, it correctly caught both AI paragraphs I gave it.
Can AI detectors give false positives on human writing?
Yes. Originality.ai flagged my own 2020 article, written years before ChatGPT existed, as 50 to 56% likely AI-generated. Undetectable AI flagged the same article as 99% AI.
Should I trust an AI detector's accuracy score?
Not on its own. The same 13-tool test that found one accurate detector also found two that flagged genuine human writing as AI. A single tool's score is a coin flip until you know which tool it is and how it's performed against known samples.
Do AI detectors get more accurate over time?
Not predictably. BrandWell missed both AI samples in my original test, then caught both correctly a year later on a fresh set. That could be a genuinely improved model or just an easier sample set, and there's no way to tell from the outside.
Want this built for your business? Book a free strategy call →