this post was submitted on 17 Aug 2026
876 points (99.4% liked)
me_irl
8185 readers
813 users here now
All posts need to have the same title: me_irl it is allowed to use an emoji instead of the underscore _
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
That’s a bold claim with no proof.
Worst performance for GPTZero is technical medical papers with a 10% false positive rate. Here’s the study from S. Korea.
Scientific scholarly writing: <4% false positive.
Creative writing is orders of magnitude easier to detect. The biggest problem for AI writing detection tools is the use of translation — so generating text, translating it into another language, and translating it back — but this ends up as a false-negative issue. False positives are very rare.
More importantly, we don’t need GPTZero to detect AI writing because professionals can identify it on sight. Add context clues (students are dumb) and suddenly it’s trivial.
EDIT: I need you to understand that college students are terrible writers. In all my decades and literally thousands of students, I’ve met maybe three undergrads whose writing was more than proficient, and that's in an environment where the average SAT scores are over 1550. The idea I’m going to mistake your writing for AI (which sounds horrible and robotic, btw) because of its “high quality” is simply unfathomable.
Thank you for linking me to that study! Unfortunately google is a piece of shit and they’re calling me a bot and refusing to talk to me even though I’m on a standard residential internet connection. Could you tell me the title/authors so I can look for it elsewhere?
I’m particularly interested in seeing what the actual false positive rates are beyond a colloquial “1%”. Because 1% is, in the context of accusing academic cheating, absolutely unacceptable dog shit.
Take a basic example intro class with 350 students, with only two tests per semester. With “only” a 1% false positive rates you will be destroying the lives of more than a dozen innocent kids every school year. And thats only for one class, students are taking many classes per semester.
Well, an important detail is that, as your professor I would already suspect you cheated. The only reason to use such a tool is to make it harder to deny.
Imagine an Ivy League classroom of 100 students, as you said. First of all, they all suck at writing in their own quirky ways. By contrast, AI has no rhythm, clumsiness, or voice. It writes like a corporation. If you’re being accused of using AI, that’s not a compliment.
Anyways, here’s another thing about these studies. Counterintuitively, a 1% false positive for a tool like GPTZero doesn’t mean that if I feed it 100 human essays from my creative writing class it’ll get one of them wrong.
It’s more like if there are 100 “types” of creative writing, 1 of them would be wrong consistently and the other 99 almost never. For these tools, there are types of writing they can’t ever figure out, like blind spots (probably from lack of data). A college essay is not going to be such a blind spot.
Here’s a known blind spot: if you have two humans collaborating on a paper, GPTZero will fail. So one person writes, the other extensively re-writes — like a translation. You end up with an inauthentic mess that throws up false positives.
If you’re looking for studies, go to Google scholar and type “GPTZero accuracy specificity” or something make sure you set it to 2025 or later, since GPTZero is rapidly improving. Most studies don’t target creative writing though, because like I said it basically doesn’t get those wrong and scientific articles are more difficult.
I’m glad I never had you as a professor, your overconfident arrogance is terrifying.
All the experts I’ve known didn’t need to tell me. It was just obvious.
Props for responding constructively.
The false positive rate is pretty low, but a false negative of 35% is pretty concerning in terms of whether they're fit for purpose. Like it'll rarely misclassify human writing as AI generated, but it's somewhat closer to a coin flip in correctly identifying AI generated language. It's a signal, but I imagine it's a poor substitute for the judgement of a person who has spent time teaching teenagers / young adults to write.
Which to the original point isn't totally snake oil, but it's maybe not a great tool for the professionals who it's targeting.
The false negative rate for creative writing is exceedingly low as well, sub 1%. False negatives aren’t the critique of this meme, and the examples I gave for GPTZero were worst case scenarios for very specific types of technical writing.