Introduction
Do you know somebody who has been falsely accused of using AI for their school or work project? Are you a manager in charge of written content and are wondering whether the texts you receive are AI-generated? Or are you merely curious about the accuracy of AI detectors? AI detectors are online tools created to detect whether a text was written by generative AI or humans.
I have been wondering for a while how accurately these AI detectors actually detect human vs. AI-generated content. So I decided to do a semi-scientific study to answer this question. The study is “semi-scientific,” because instead of testing these AI detectors on a large scale with randomly selected texts to obtain statistically meaningful results, I chose a handful of manually curated texts. However, the rest of the methodology is as scientific as possible, with detailed, auditable records of all results. More on the methodology below.
The question I set out to investigate qualitatively (but not quantitatively) is:
How common are false positives when AI detectors are tasked to determine if a text was AI-generated?
Now, for the less mathematically inclined, the term “false positive” warrants an explanation. The rest of the readership may skip this paragraph. The question that AI detectors investigate is: “Is it AI?” If the answer is affirmative, then that is called a positive. Therefore, if the text in question was generated by AI, and the AI detector says so, then it’s a true positive. But if the text in question was instead generated by a human, but the AI detector says it’s AI, then the result is false, that is, a false positive. In other words, the question I wanted to study was how often 100% human-generated text is misidentified as AI by these AI detectors.
Methodology
Text Selection
The texts were selected with the following constraints. First, I needed 100% verifiably human texts not constrained by copyright or other legal issues. Then, the text passages needed to be longer than 200 words, but preferably less than 750 words (or 1000 words). The lower limit is because virtually all available AI detectors state that they cannot reliably detect AI usage for shorter passages. The upper limit stems from the fact that I wanted to peruse the free trials for the chosen AI detectors, since this study is out of mere curiosity and not part of a formal (funded) research program.
Now, where was I going to find 100% verifiably human, scientific text passages? Simple, I chose older texts. The reason for selecting mostly my own writing is not to show off my research career, but rather to avoid any copyright issues. I chose the following sources:
- My own PhD thesis from May 2003, guaranteed AI-free and verified by my PhD thesis committee. The thesis is freely available for download from the arXiv here.
- My most-cited physics publication, which I co-authored in 2008: An Automated Implementation of On-Shell Methods for One-Loop Amplitudes, C. F. Berger et. al, Phys. Rev. D 78 (2008) 036003, the preprint is available for download here.
- To change things up, a patent on bicycles by George Seyfang from 1903, where the patent protection expired long ago, US Patent No. 730,194, available for download here.
From these sources, I picked six passages that fulfilled the aforementioned length requirements and saved them as plain text files, which I then copy-pasted into the online AI detectors.
For good measure, I added one AI-generated text to the mix. This text was generated by the free version of ChatGPT on July 1st, 2025. I prompted ChatGPT to create a 500 word summary of my publication in plain text format, Higher orders in A(alpha(s))/[1-x]+ of nonsinglet partonic splitting functions, C. F. Berger, Phys. Rev. D 66 (2002) 116002, the preprint is available for download here. I did not edit the output at all, not even for factual errors.
All passages are listed below with their pertinent details, accordion-style, so as to not take too much screen space.
AI Detector Selection
I picked four well-known AI detectors that offered a free trial without signing in, or at most with signing in with a Google account. I also picked two AI detection aggregators that purported to aggregate the results of several individual AI detectors, including the four individual engines I tested. My choice of AI detectors is merely a random sample of all detectors out there. Beyond that, you should not infer any other reason for selecting these and leaving out others.
The detectors I tested are:
- GPTZero, model 3.5b, and earlier model 3.4b
- Grammarly
- Originality.ai
- QuillBot, v.5.3.1
- Aggregator undetectable AI, states that it uses GPTZero, OpenAI, Writer, QuillBot, Copyleaks, Sapling, Grammarly, ZeroGPT
- Aggregator JustDone, states that it uses Originality.ai, Scribbr, GPTZero, Copyleaks
GPTZero and QuillBot state a model/version number, the rest does not, which is why I dated the tests. As we will see below, the results evolve over time and depend on the specific model version.
Running the Tests
The rest is straightforward: I simply took each of the text passages above and fed it through each of the AI detectors. For reference and auditability, I took screenshots of all results. I also entered the results into a simple spreadsheet. Obviously, the number of tests is by far not sufficient to come to any statistically relevant conclusions, but some of the qualitative results are quite intriguing. Most of the tests were run on June 30, 2025 and July 1, 2025, but a couple were done on May 30, 2025, when I first started this study, but then got sidetracked with paying projects. This timeline led to some unexpected and interesting secondary results, which I will now present along with the main findings.
The Results
First Attempt – PhD Thesis Abstract
My very first attempt at the end of May was to simply grab the first 100% human text I could find and run it through the aforementioned AI detectors. This first text, Text number 1 above, happened to be the abstract of my PhD thesis, written in 2003, at a time, when the only AI people knew about was a township in Ohio (it still is). All four individual detectors, GPTZero, Grammarly, Originality, and QuillBot correctly identified the text as 0% AI. Interestingly, JustDone claimed that the text was 93% AI (as of June 30, 2025), although JustDone says that it uses Originality and GPTZero to “double-check” the result, as can be seen in the screenshot below.

Even more interestingly, I ran the same abstract through undetectable AI twice, once on May 30, 2025, and once on June 30, 2025. On May 30, 2025, it identified the text as 61% AI. Then, notably, on June 30, 2025, undetectable AI gave a result of 1% AI. undetectable AI does not state a version/model number, but the results clearly vary over time.


Results for 100% Human Texts
Grammarly and originality.ai all correctly output 0% AI for all human-written texts, regardless of length, and undetectable AI output 1% AI, which is essentially equivalent to 0%. Interestingly, QuillBot version 5.3.1, claimed that 11% and 5%, of Text 4 and Text 5, while written by humans, were refined with the help of AI, as can be seen in the screenshots below. I would like to reiterate that generative AI engines did not exist at the time Texts 4 and 5 were written, in 2008, so they could not have been refined by any artificial method at all, except perhaps a spell checker.

GPTZero model 3.5b correctly gave 0% AI for all human-written texts, but interestingly, GPTZero model 3.4b, available in May 2025, claimed that 7% of Text 5 was again refined with the help of AI.

JustDone’s output, on the other hand, claimed that a large percentage of each of the human-written texts was written by AI, namely as follows:
- 93% for Text 1
- 75% for Text 2
- 83% for Text 3
- 77% for Text 4
- 84% for Text 5
- 74% for Text 6
Results for the AI-Generated Text
Finally, I sent the entirely AI-generated summary of my old paper through the AI detectors. The results were a mixed bag. Grammarly and QuillBot v. 5.3.1 both claimed the text was entirely human (0% AI), and both originality AI and undetectable AI concurred with that statement at 1% AI. Only GPTZero model 3.5b output that 31% of the text was generated by AI and another 11% was refined by AI, whereas the remaining 52% were allegedly written by a human.

Interestingly, JustDone output that 72% of the AI text was generated by AI, its lowest percentage of all texts that I tested. By contrast, the hand-written texts above scored a higher AI percentage than the actually AI-generated text. In other words, JustDone identified the AI text as the most human of the seven texts that I tested.
Conclusion: Caveat Emptor
While I did not have the time, money, or energy to run a statistically significant amount of human and AI-generated texts through these AI detectors to come to a quantitative conclusion, it is qualitatively safe to say that the results are certainly a mixed bag. Sometimes human text was correctly identified as such, sometimes not, and sometimes AI-generated text was identified as human, and sometimes as partially AI.
AI detectors are nothing but generative AI engines (also known as Large Language Models, LLMs) finetuned to detect AI output. And as such, AI detectors seem to be just like any other AI tool out there—sometimes they give good results, but more often than not, the results need to be double-checked by thinking humans instead of blindly believing the AI output. Further, the results seem highly model-dependent and thus evolve over time along with the evolution of the underlying language model. The results I presented above from May-July 2025 may not hold in a year or even a month from now.
Where does that leave us? I believe that it’s the content, not the writing style, which will be the distinguishing feature between AI-generated content and human pieces. AI is (currently) incapable of original thought, and those humans who think before they write (or translate) will stand out, in the future even more so than now.




One Response
Hi Carola,
I’m currently watching your session for the “AI in Translation Summit”.
Thank you for checking on this important topic! I will certainly take some time later on to read this blog post carefully :-)
Liebe Grüße aus Berlin
Carolin