Freddie deBoer exposes a dangerous flaw in the digital tools currently being used to police professional writing: they are not just inaccurate, they are logically incoherent. In an era where a single false flag can end a career, deBoer demonstrates that the very software marketed as the solution to AI cheating is capable of declaring a 5,000-word essay "100% human" while simultaneously flagging a 300-word excerpt within it as "100% AI." This isn't a minor glitch; it is a fundamental breakdown of the instrument's reliability that threatens to upend how we evaluate truth and authorship.
The Paradox of the Percentage Meter
The core of deBoer's argument is that the "percentage meter" displayed by detection tools like Pangram is not measuring what users think it is. Many assume a "100% AI" reading means the system is 100% confident the text is machine-generated. DeBoer clarifies that the percentage is supposed to indicate the portion of the text suspected of being AI-written, a distinction that is crucial but frequently ignored. The tool's output, however, defies basic logic. "The percentage meter seems clearly broken," deBoer writes, noting that he can take a section flagged as AI, break it into smaller pieces, and have those pieces flagged as human. Conversely, he can embed human text into AI writing and have the whole thing flagged as AI.
This inconsistency creates a "Russian doll of contradictory results," where the whole is human, the part is AI, and the sub-part is human again. As deBoer puts it, "These claims are mutually exclusive and must necessarily erode our confidence in the instrument." The tool's design forces users to choose between conflicting data points, effectively turning the detector into a Rorschach test where observers see what they want to see. This is particularly alarming given the stakes; these tools are being deployed in high-pressure environments where a false positive can lead to termination or academic expulsion.
"The percentage is supposed to indicate what portion of the text is suspected to be LLM written. The degree of confidence is flagged there underneath 'AI Generated,' although annoyingly the flag only appears when confidence is high, with no 'Confidence Low' flag that ever appears, in my experience."
DeBoer's critique here is sharp because it targets the user interface's failure to communicate uncertainty. In statistics, a lack of a "low confidence" flag is a significant design flaw, as it forces a binary interpretation on a probabilistic result. This mirrors historical issues in stylometry, where early attempts to attribute authorship often failed because they could not account for the natural variance in a single writer's voice across different contexts. Just as Goodhart's law warns that "when a measure becomes a target, it ceases to be a good measure," the pressure to get a definitive "100%" result has likely warped the algorithm's calibration, pushing it toward extreme, polarized outputs rather than nuanced probabilities.
The Mathematics of False Accusations
The most terrifying aspect of deBoer's analysis is the statistical inevitability of false positives when these tools are applied at scale. Even if a tool has a low error rate, the sheer volume of text being analyzed guarantees that innocent writers will be caught in the net. DeBoer calculates that with his own output of roughly 1,200 posts since 2022, and the detector operating on a sentence or paragraph level, he faces approximately 18,000 opportunities to be falsely flagged. "Even a very small error rate will produce false outcomes regularly given enough repetitions," he notes, arguing that a "one drop rule"—where a single flagged sentence condemns the whole work—is mathematically unsound.
This is where the human cost of the technology becomes undeniable. DeBoer warns that "some people's natural writing style is more likely to be flagged than that of others, even with no LLM assistance." He illustrates this with the example of non-native English speakers, suggesting that linguistic patterns common among international students could be misidentified as AI artifacts. "Can't you imagine the outrage if it's eventually found that, say, Chinese international students are more likely to produce false positives with their writing?" he asks. This is not a hypothetical; it is a direct consequence of training models on data that may not represent the full diversity of human expression.
Critics might argue that in an age of rapid AI proliferation, we must accept some level of false positives to catch the bad actors. However, deBoer's demonstration that the tool is easily manipulated by context undermines this defense. If a human writer can be tricked into a "100% AI" verdict simply by adding a few sentences of AI text, or if the tool's verdict changes based on the surrounding paragraph, then the tool is not a reliable arbiter of truth. As deBoer writes, "I can take AI-generated text and embed it in human-generated text and have Pangram declare it 100% human, and I can take human-generated text and embed it in AI-generated text and have Pangram declare it 100% AI." This brittleness suggests the tool is measuring something other than authorship—perhaps a specific, narrow set of stylistic markers that can be easily gamed or that happen to overlap with legitimate human writing styles.
The Danger of Confirmation Bias
The ultimate danger, deBoer argues, is not just the technical failure of the software, but the social dynamic it encourages. When the data is contradictory, people will not seek the truth; they will simply select the data point that confirms their existing biases. "This is what I worry about with this whole controversy - we're very likely to end up with everybody confirming their priors," deBoer writes. Those who dislike a writer will point to the flagged excerpt as proof of cheating, while supporters will point to the full-essay clearance as proof of innocence. The result is a stalemate where no one is convinced, and the tool serves only to deepen divisions.
"Nobody will be convinced of anything. The nightmarish possibility is that people will start cutting up essays that pass Pangram's test into small pieces, looking to find parts that get flagged as AI, so that nobody ever feels confident about anything."
This fragmentation of evidence is a recipe for chaos. It transforms the act of writing into a minefield where every sentence could be weaponized against its author. DeBoer's call for a more rigorous, corpus-based approach—looking at a writer's entire body of work rather than isolated snippets—is a necessary corrective. However, he admits this is "hard and time intensive and, most importantly, doesn't fit with the desire to own one's enemies online." The tool is not being used for genuine detection; it is being used as a "one-shot gotcha machine" to settle scores.
Bottom Line
Freddie deBoer's dissection of Pangram reveals a tool that is not merely imperfect, but fundamentally broken, offering contradictory results that undermine its own purpose. While the argument that statistical error rates will inevitably victimize non-native speakers and consistent writers is compelling, the piece's greatest strength is its demonstration of how easily the tool can be manipulated by context. The biggest vulnerability in the current landscape is the willingness of institutions to rely on these brittle metrics for life-altering decisions without understanding the underlying mathematics of false positives. Readers should watch for how this technology evolves, but more importantly, they should demand that any system used to judge human authorship be held to a standard of logical consistency that Pangram currently fails to meet.