A single line of inverted code has the power to unravel a decade of economic history, yet the prestigious journals that published the flawed findings refuse to admit the mistake. Brad DeLong uses this specific technical failure to expose a rot in the academic incentive structure, arguing that the "top five" economics journals have become so obsessed with their own infallibility that they ignore obvious errors. This isn't just about a math mistake; it is a case study in how confirmation bias and institutional arrogance can cement false narratives in the public record for years.
The Mechanics of a Broken Finding
DeLong introduces the work of Joe Francis, a researcher who performed the public service of auditing a highly cited paper by Oded Galor and Ömer Özak. The paper, published in the American Economic Review, claimed that historical crop yields shaped modern cultural patience. DeLong highlights the sheer scale of the error: "A tiny 34 km² Swiss region receives 170,000 times the weight of the largest region." This inversion meant the study was effectively amplifying the noise of small, urban areas where migration is highest, rather than dampening it as the authors claimed.
The core of the argument rests on a psychological trap familiar to anyone who has analyzed data. As DeLong writes, "You have a prior—a strong, theoretically motivated, Malthusian-selection prior that patience got culturally bred into populations where the agronomic return to waiting was high." When the results matched this deep-seated belief, the authors stopped auditing the code. DeLong notes, "Confirmation is where your guard is down." This is a critical insight that transcends economics; it mirrors the replication crisis seen across the social sciences, where the pressure to find a signal often overrides the discipline of checking the machinery.
"The estimator gives the largest weight to the smallest regions... The influence of small, urban regions where internal migration is most severe is thereby maximized, making the problem that area weights were supposed to address worse."
When the code was corrected, the entire story collapsed. The statistical significance vanished, and the coefficient even flipped signs. DeLong points out that this is not a rare anomaly but a symptom of a system where "people have moved" and macro-historical regressions are "famously vulnerable to confounding factors." The study tried to patch this with a half-hearted weighting scheme, but the inversion turned a patch into a wound.
The Institutional Silence
The most disturbing part of DeLong's commentary is not the error itself, but the reaction—or lack thereof—from the academic establishment. He cites a grim statistic from The Economist to illustrate the rigidity of the field: "the five leading journals have seen just four withdrawals in their combined 570-year history." DeLong argues that this creates a perverse dynamic where the peer review process is so mythologized that admitting a mistake is seen as impossible.
Francis, the original author of the critique, expresses a "counsel of despair," noting that "even if I find a fault in one of the articles published in the 'top five' economics journals, I know that it will not be retracted or even corrected." DeLong pushes back against this hopelessness, suggesting that while formal retractions are rare, the "machinery works, albeit maddeningly slowly." He argues that the profession eventually weeds out non-replicable results, even if it takes a decade and a dedicated outsider like Francis to do the work.
Critics might note that relying on a slow, informal correction process is an inadequate defense for a field that claims scientific rigor. If a finding is fundamentally broken, the fact that it eventually stops anchoring dissertations does not excuse the years of policy or academic work built on its back. DeLong acknowledges this tension, admitting, "It is a shame it takes a decade and a Joe Francis rather than an afternoon and a journal's own referees."
The Path Forward
DeLong concludes by shifting the focus from the specific error to the broader methodological failures of "deep roots" literature. He argues that these studies often rely on cross-sectional correlations to masquerade as causal parameters without a structural model to explain the mechanism. "I do not believe you have earned the right to say 'causal' unless you can write down the structural model," DeLong asserts, positioning himself as a "Heckmanite" who demands more rigorous proof of exogeneity.
The solution, according to DeLong, lies in structural changes to how research is conducted and published. "Which is why we need preregistration and replication packages," he writes, framing Francis's work not as a complaint, but as a necessary evolution of the field. The inversion of the area weights was a coding slip, but the refusal to correct it is a cultural one. As DeLong puts it, "Things that do not replicate are weeded out, albeit much more slowly than they should be."
"It is not just computational error. It is also that, say, you have ten yes-no decisions to make in running any empirical analysis that could go either way... Succumb to that, and you wind up with the strongest of 1024 possible results."
Bottom Line
DeLong's commentary is a powerful indictment of an academic culture that prioritizes narrative consistency over empirical truth, using a specific coding error to illustrate a systemic failure of self-correction. While his defense of the profession's eventual ability to self-correct is optimistic, it risks underestimating the real-world damage caused by the decade-long lag between error and correction. The strongest takeaway is the urgent need for preregistration and a cultural shift that treats replication not as an attack, but as the essential engine of scientific progress.