- 0 Posts
- 7 Comments
Well, an important detail is that, as your professor I would already suspect you cheated. The only reason to use such a tool is to make it harder to deny.
Imagine an Ivy League classroom of 100 students, as you said. First of all, they all suck at writing in their own quirky ways. By contrast, AI has no rhythm, clumsiness, or voice. It writes like a corporation. If you’re being accused of using AI, that’s not a compliment.
Anyways, here’s another thing about these studies. Counterintuitively, a 1% false positive for a tool like GPTZero doesn’t mean that if I feed it 100 human essays from my creative writing class it’ll get one of them wrong.
It’s more like if there are 100 “types” of creative writing, 1 of them would be wrong consistently and the other 99 almost never. For these tools, there are types of writing they can’t ever figure out, like blind spots (probably from lack of data). A college essay is not going to be such a blind spot.
Here’s a known blind spot: if you have two humans collaborating on a paper, GPTZero will fail. So one person writes, the other extensively re-writes — like a translation. You end up with an inauthentic mess that throws up false positives.
If you’re looking for studies, go to Google scholar and type “GPTZero accuracy specificity” or something make sure you set it to 2025 or later, since GPTZero is rapidly improving. Most studies don’t target creative writing though, because like I said it basically doesn’t get those wrong and scientific articles are more difficult.
The false negative rate for creative writing is exceedingly low as well, sub 1%. False negatives aren’t the critique of this meme, and the examples I gave for GPTZero were worst case scenarios for very specific types of technical writing.
Worst performance for GPTZero is technical medical papers with a 10% false positive rate. Here’s the study from S. Korea.
Scientific scholarly writing: <4% false positive.
Creative writing is orders of magnitude easier to detect. The biggest problem for AI writing detection tools is the use of translation — so generating text, translating it into another language, and translating it back — but this ends up as a false-negative issue. False positives are very rare.
More importantly, we don’t need GPTZero to detect AI writing because professionals can identify it on sight. Add context clues (students are dumb) and suddenly it’s trivial.
EDIT: I need you to understand that college students are terrible writers. In all my decades and literally thousands of students, I’ve met maybe three undergrads whose writing was more than proficient, and that’s in an environment where the average SAT scores are over 1550. The idea I’m going to mistake your writing for AI (which sounds horrible and robotic, btw) because of its “high quality” is simply unfathomable.
Nobody who reads or writes for a living thinks AI slop sounds good. It has a hideous corpo-speak prosody that’s like nails on a chalkboard to professionals.
People think because they can’t tell it’s not written by a human, someone who has read literally tens of thousands of essays (not to mention millions of fucking book pages) can’t tell either.
Absolutely bonkers.
Students who use AI are like little kids standing on each others shoulders in a trench coat and trying to order alcohol.
It’s actually the opposite. AI writing sounds like corpo-speak. It’s soulless, ugly, and extremely obvious.
I’ve spent about 30,000 hours reading in my life. If you think I can’t tell you’re using a token predictor for your “essay,” you’re fucking insane.
They’re not. GPTZero has a false positive rate under 1% for creative writing.

He’s being sarcastic.
deleted by creator