← Back to Library

Model collapse

Cory Doctorow doesn't just warn about AI; he diagnoses a systemic rot where the very tools meant to predict our future are actively erasing it. By weaving together urban planning, music theory, and economic policy, Doctorow reveals a terrifying paradox: the more data we feed our algorithms to "personalize" our world, the more homogenized and bland that world becomes. This isn't just a tech glitch; it is a cultural and economic feedback loop that threatens to flatten human diversity into a single, safe, and profitable average.

The Centaur Paradox and the Investment Bubble

Doctorow begins by dismantling the confusing narrative around AI's capabilities. He argues that the conflicting reports of AI success and failure stem from a fundamental misunderstanding of the human-machine relationship. "One of my favorite rhetorical and analytical moves is joining things together... and taking them apart," he writes, applying this to the AI workforce. He distinguishes between "centaurs" (humans assisted by machines) and "reverse centaurs" (humans forced to act as peripherals for machines). This distinction is crucial because it shifts the blame from the technology's inherent flaws to the management decisions that prioritize cost-cutting over competence.

Model collapse

The author suggests that the current AI boom is fueled not by a belief in the technology's actual utility, but by a cynical assessment of sales potential. He notes that investors are often betting on the ability to sell AI to credulous bosses who want to replace difficult workers with pliable software. "They don't have to believe in AI in order to think it's a good investment: like an investor betting that Joe Rogan can sell millions of dollars' worth of peptides to desperate young men, they are assessing the sales potential, not the merits of the thing for sale." This framing exposes the hollow core of the current market frenzy, suggesting that the bubble is sustained by a willingness to ignore the fact that the product often cannot do the job it was hired for. Critics might argue that this view is overly cynical and ignores genuine productivity gains in specific, narrow domains, but Doctorow's point holds weight when looking at the broader trend of mass layoffs followed by a failure to meet output targets.

"Personalisation under a standard loss function is regression to the collective mean with extra steps."

The Self-Fulfilling Prophecy of Data

The core of Doctorow's argument relies heavily on the work of data scientist Lauren Leek, whose essay "Temperature Zero for Culture" serves as the intellectual backbone of the piece. Doctorow synthesizes Leek's findings to show that predictions in a data-driven society are not passive observations; they are active forces that reshape reality. "Once prediction shapes the choices in front of us, we lose the ability to tell the difference between what people wanted and what the system made easy to want," Doctorow quotes. This concept, known as "performativity" in economics, explains why markets and cultures begin to look identical.

He connects this to the historical concept of Goodhart's Law, which states that "When a measure becomes a target, it ceases to be a good measure." Doctorow illustrates this with the example of Google's PageRank. Originally, counting inbound links was a brilliant way to find quality content. But once that metric became the target, it was gamed, and eventually, it stopped predicting quality and started predicting who could game the system best. Similarly, in the physical world, this leads to "placelessness." Doctorow describes this as "Flinstones Syndrome," where every street looks the same because algorithms optimize for the same variables. "A relatively small number of repeated names is enough to make otherwise different streets resemble one another more," he notes, pointing to how chain stores use footfall data to ensure every location feels identical to every other location.

This dynamic is devastating for local culture. Doctorow highlights Leek's research on UK pub closures, which found that the biggest predictor of a pub's survival was its similarity to the median pub. "The more distinctive a pub was, the more 'character' it had, the more likely it was to close," he explains. Algorithms used by banks and landlords cannot assess the value of uniqueness, so they starve it of capital. This mirrors the tragedy of the commons, where individual optimization leads to collective ruin, but here, the ruin is cultural blandness.

The Collapse of Culture

The most chilling application of this theory is in the realm of media and art. Doctorow explains that music recommendation systems, optimized for "singable hooks," have forced songs to use a smaller vocabulary of unique words and repeat them more often. "Vocabulary richness, distinct words relative to length, has fallen by more than a quarter since the early 1960s," he writes. While there is a superficial variety in the types of words used, the structure of the music has flattened. The algorithms are not discovering new tastes; they are narrowing the menu.

He cites a stark statistic from Movietweetings: out of a million public movie ratings, half relate to the top 2% of movies. "There's 38,000 films in the set, but just 380 titles account for 40% of the ratings," Doctorow points out. This isn't because audiences have suddenly converged on the same taste; it's because the recommender systems have "narrowed the menu" so effectively that the long tail of diverse options is never even presented to the user. The features of media that might have been appreciated are "decaying out of consideration" because they were never tested for desirability.

This leads to the phenomenon of "model collapse," where AI models trained on their own outputs become increasingly bland. Doctorow references Leek's failed attempt to create synthetic personas for market research. Despite painstakingly replicating demographic and psychological factors, the resulting AI personas were "homogenized average" versions of humans. "It's like the paradox of 'The Average Man,' where military uniforms sized to the average of all service personnel fit no one, because no one is average," he writes. The danger is that these synthetic personas are now being sold to governments and politicians to model public opinion, creating a feedback loop where policy is designed for a fictional, average population that doesn't exist.

"They don't have to be right, only listened to."

Bottom Line

Doctorow's most powerful contribution here is his synthesis of disparate fields to prove that "model collapse" is not a future risk but a present reality affecting our streets, our songs, and our laws. The argument's greatest strength is its refusal to treat AI as an isolated technological issue, instead showing it as the accelerant for a broader crisis of standardization. However, the piece could be strengthened by addressing how specific regulatory interventions might break this feedback loop, rather than just diagnosing the problem. As we move forward, the critical question is not whether AI will replace us, but whether we will allow it to replace the very diversity that makes us human.

Deep Dives

Explore these related deep dives:

  • Goodhart's law

    The article uses this principle to explain how AI metrics become gamed when optimization targets replace genuine quality, leading to the degradation of output.

  • Gros Michel

    This extinct banana variety serves as a historical analogy for how monocultures and homogenized systems collapse when they lose the diversity needed to adapt to new threats.

  • Tragedy of the commons

    This concept illuminates the article's argument that training AI on AI-generated content is a form of digital enclosure where the shared pool of human creativity is exhausted and polluted by its own synthetic byproducts.

Sources

Model collapse

by Cory Doctorow · Pluralistic · Read full article

Today's links.

Model collapse: Living in a world that's trained on itself. Hey look at this: Delights to delectate. Object permanence: Wired v Dutch hackers; NYT v DMCA; Hair gel terrorist threat does not exist; AT&T merger is a screwjob; Smart cities are stupid; RIP Reaganomics. Upcoming appearances: Edinburgh, Sydney, Melbourne, Brighton, London, South Bend. Recent appearances: Where I've been. Latest books: You keep readin' em, I'll keep writin' 'em. Upcoming books: Like I said, I'll keep writin' 'em. Colophon: All the rest.

Model collapse (permalink).

One of my favorite rhetorical and analytical moves is joining things together (showing that two different, seemingly unrelated ideas are aspects of the same phenomenon) and taking them apart (resolving a paradox by demonstrating that what appears to be one, contradictory thing is actually two different things that have been lumped together).

"Taking things apart" is a very useful framework for understanding AI. How do we resolve the (seeming) paradox that some skilled workers report wonderful results from their work with AI, while others are full of dire warnings about the lurking defects in their AI-assisted outputs? Simple: the first group are "centaurs" (humans who are assisted by machines) and the second are "reverse centaurs" (humans who have been pressed into service as peripherals for machines):

https://pluralistic.net/2025/12/05/pop-that-bubble/#u-washington

What are we to make of the people who've been fired by bosses who replaced them with AI, in light of the fact that AI is demonstrably not able to do their (former) jobs? Again, it's simple if you separate out two distinct phenomena: "AI can do your job" is the first. The second is: "Your boss is a credulous dolt who is infinitely horny for replacing lippy workers with pliable machines, which made him an easy mark for an AI salesman who convinced him to fire you and replace you with an AI that can't do your job":

https://pluralistic.net/2025/03/18/asbestos-in-the-walls/#government-by-spicy-autocomplete

This is also a useful move for understanding the AI investment bubble. It's not just billionaires who don't think other people are as real as they are and consequently their jobs can be done by chatbots. It's also billionaires who believe that bosses can be sold AI and don't care if the AI is defective, because that's your boss's problem after he buys the AI and fires you. They don't have to believe in AI in order to think it's a good investment: like an investor betting ...