In an era where digital noise threatens to drown out human voice, Alberto Romero makes a startlingly pragmatic case: we must accept the occasional wrongful accusation to save the integrity of the written word. While most critics decry the new Substack-Pangram partnership as a slippery slope toward censorship, Romero reframes it as a necessary, albeit imperfect, defense against a "tragedy of the commons" where low-effort AI generation could render the internet unreadable.
The Lesser Evil of Digital Witch Hunts
Romero does not shy away from the uncomfortable historical parallels. He acknowledges that the partnership between the publishing platform and the AI detection firm feels like a "witch hunt," yet he argues that the context justifies the method. "If back in the seventeenth century in Salem or Scotland, witches had actually been killing people with potions, ointments, and sorcery, then hunting them would've been strictly a good thing," Romero writes. He posits that while the tool is blunt, the threat of "AI slop"—content generated en masse to game algorithms rather than inform readers—is a genuine existential threat to the platform's value.
This framing is provocative because it prioritizes the health of the ecosystem over the absolute rights of the individual user. Romero admits that detection systems are flawed, noting that "every detection system is imperfect." However, he argues that the alternative—a web flooded with content that no human actually wrote—is a far worse outcome. He accepts the "tragedy of the unlucky," where innocent writers might be flagged, as the price for a cleaner digital environment. This utilitarian calculus is effective because it strips away the moral panic and focuses on the practical reality of scale: when bad actors flood the zone, perfection in enforcement becomes impossible.
Critics might note that this logic echoes the very authoritarian tendencies it claims to fight, where the definition of "slop" is subjective and the punishment for being flagged can be social and financial, regardless of the technical error rate.
"We're not casualties of a platform-wide cleansing but martyrs for a digital golden age."
The Mechanics of Stylometry and the False Positive Trade-off
Moving beyond the philosophy, Romero dives into the technical specifics of Pangram, the detection tool at the heart of the controversy. He highlights the company's aggressive focus on minimizing false positives—instances where human writing is incorrectly flagged as AI. Romero notes that Pangram has pushed this rate down to a theoretical minimum of 0.0041%, or roughly one error for every 24,000 instances. "Pangram is not perfect, but it's close to being as good as a detector can be," he observes.
The author's argument here is grounded in the concept of "stylometry," the idea that AI models, despite their ability to mimic human prose, leave behind distinct "stylistic fingerprints" or a "styleme." Romero recalls predicting in late 2022 that AI would develop its own idiosyncrasies, just as human authors have unique voices. He argues that the tool doesn't rely on naive heuristics but rather on the aggregate conditioning of the model itself. "You, a mere human, can tell AI writing apart from human writing if you train enough," he asserts, suggesting that the technology is simply scaling up a capability that humans already possess.
However, Romero is careful to point out the inherent trade-off in this technology. To ensure that the few people flagged are almost certainly guilty, the system must accept a higher rate of false negatives—letting some AI content slip through. He quotes the 18th-century legal philosopher William Blackstone: "it's better for 10 guilty people to escape than for for 1 innocent person to suffer." Romero admits this principle is being bent here, but argues that the specific context of AI-generated spam necessitates a different balance. The system is designed so that "everything flagged as AI will almost always be AI," making the accusation a powerful deterrent even if it isn't a perfect filter.
Deterrence Over Punishment
Perhaps the most crucial distinction Romero makes is about the actual mechanism of enforcement. He clarifies that Substack is not using Pangram to ban users or issue lifetime suspensions. "If you're flagged as AI, it doesn't result in a lifetime ban from the platform," he writes. Instead, the tool serves as an "externally enforceable disclosure," giving readers the information they need to decide whether to engage with a piece of content.
This reframes the partnership from a punitive measure to an informational one. Romero argues that the goal is to create a market where readers can avoid "AI slop," thereby disincentivizing its creation. "The goal behind this partnership is, instead, to give people the ability to know what's going on," he explains. If successful, this could create a ripple effect, forcing universities, newsrooms, and other platforms to adopt similar transparency measures. He envisions a future where the "rest of the world will have no choice but to follow suit, for there will be no excuses whatsoever not to kill slop."
This approach attempts to solve the "paradox of tolerance"—the idea that a tolerant society must be intolerant of intolerance to survive. Romero suggests that a society that tolerates the flood of AI-generated content is effectively tolerating its own obsolescence. "Our society is at stake—our humanity is at stake," he warns, borrowing Lincoln's words to emphasize the gravity of the moment. He argues that the risk of encountering fake content is so high that it would drive newcomers away, much like a moldy shower plate would repel a guest.
The Danger of the Mob
Despite his support for the initiative, Romero issues a stern warning about the human element: the "mob." He argues that while the tool is precise, the people wielding it may not be. "A witch hunt, even when witches exist and non-witches are mostly safe, is still governed by the mob. And the mob, more so than any other agent involved here—including AI—is stupid," he writes. He urges readers not to use the detection scores as a substitute for their own judgment or to engage in public shaming without understanding the tool's limitations.
He points out that Pangram cannot distinguish between a writer who used AI for brainstorming versus one who outsourced the entire essay. "Substack was careful about this," Romero notes, quoting their disclaimer that the tool detects AI usage but not the level of human care involved. Yet, he fears that "the mob is the enemy of nuance" and will make absolute accusations based on relative data. He challenges the community to be more rigorous in their understanding of the technology than the bad actors are in their use of it. "Don't go hunting without knowing the limits of the tool you wield," he commands, suggesting that the true failure would be if the community becomes as lazy in its thinking as the grifters it seeks to expose.
"Don't be the mob."
Bottom Line
Alberto Romero's argument is a bold, utilitarian defense of digital hygiene that prioritizes the survival of human-centric platforms over the absolute protection of every individual user from error. While his acceptance of false positives as a necessary evil is logically consistent within his framework, it remains the argument's most vulnerable point, relying on the hope that the community will exercise restraint and nuance rather than succumbing to the very mob mentality he warns against. The success of this initiative will depend less on the accuracy of the algorithm and more on the wisdom of the readers who wield it.