← Back to Library

Wait! What?!?!: “Mr albouy reaches his conclusion by omitting half the data from the original…

In a field obsessed with statistical significance, Brad DeLong identifies a methodological house of cards that threatens to topple one of the most famous papers in modern economics. He argues that a recent defense of the Acemoglu, Johnson, and Robinson (AJR) thesis relies not on robust data, but on a statistical artifact so severe it renders their primary conclusions mathematically meaningless. For anyone trying to understand the true drivers of global wealth, this is not just an academic squabble; it is a warning that the foundational evidence for institutional determinism may be built on a ghost.

The Methodological Collapse

DeLong opens by dismantling a recent, anonymous critique from The Economist that attempted to discredit economist David Albouy. The publication claimed Albouy reached his conclusions by "omitting half the data from the original sample," a charge DeLong labels as "simply wrong." Instead, DeLong highlights Albouy's actual, narrower, and devastating point: the original study lacks a valid "first stage" for its instrumental variables regression.

Wait! What?!?!: “Mr albouy reaches his conclusion by omitting half the data from the original…

The core of the argument rests on how the original authors handled mortality data. DeLong explains that when Albouy corrected for statistical clustering and removed 36 conjectured mortality rates—data points assigned based on guesses about similar disease environments rather than recorded history—the link between settler mortality and modern institutions "collapses toward one-in-three significance." This is a critical distinction. In econometrics, if your instrument (the tool used to prove causation) is weak, your entire model falls apart. DeLong writes, "The second-stage test statistic isn't a t-distribution; it's near-Cauchy — infinite variance, no mean. The IV estimates are unreliable."

This reference to the Cauchy distribution is not mere jargon; it is a mathematical way of saying the results are chaotic and unpredictable. To put it in historical context, this mirrors the pitfalls seen in studies of colonial origins where data gaps were filled with assumptions, a problem that has plagued comparative development research since the early attempts to quantify the impact of the British Empire. When the data is this fragile, the resulting policy implications are dangerous.

The IV estimates are highly unreliable, and the distribution of their test statistics is near to a Cauchy distribution: that thing that not only has infinite variance and standard deviation, but does not even have a mean.

Critics might argue that dropping data points introduces selection bias, but DeLong counters that keeping them introduces a far worse error: the illusion of precision where none exists. The original authors, DeLong notes, should be "very grateful" to Albouy for exposing this, yet they have instead doubled down on a defense that ignores the statistical reality.

The Implausibility of the Results

Beyond the technical failure, DeLong points out that the results themselves are absurd. If one were to accept the original paper's numbers at face value, the implied economic effects are so large they defy logic. DeLong uses a vivid metaphor to describe the magnitude of the error: "a clock that chimes thirteen."

He breaks down the math to show the disconnect. The standard correlation (OLS) suggests that better property rights lead to a 68.5% increase in prosperity. However, the flawed instrumental variable (IV) method claims a 157% increase. DeLong illustrates the absurdity: "AJR's IV results say that the effect of security-of-property on prosperity ought to be much much bigger than that: it ought to be about 20 times richer."

This creates a paradox. If the IV results were true, it would imply that prosperity actively destroys governance quality, a notion that contradicts decades of political science. DeLong lists the overwhelming evidence against this, from the Modernization Hypothesis to historical records of revolution, which show that institutional breakdown happens in poverty, not wealth. He writes, "any claim that prosperity structurally reduces governance quality contradicts an overwhelming body of evidence across multiple disciplines."

The original authors, Acemoglu, Johnson, and Robinson, defend their work by claiming their method corrects for measurement error. DeLong argues this defense is a trap. If they are right that their method isolates the "true" institutions, then the data must show that settler mortality correlates with good institutions while ignoring bad ones. But as DeLong notes, "settler mortality in the age of imperialism has to (a) be correlated with that part of modern-day institutions that matter for prosperity... while also (b) being uncorrelated with those parts of modern-day institutions that do not matter." He suggests this is an impossible standard to meet, rendering the "correction" a fabrication of the data's own noise.

In the real world, a country like New Zealand with a log prosperity score of 10 and a perceived security-of-property score of 10 was back before 2000 about 5 times richer than an Egypt... AJR's IV results say that the effect of security-of-property on prosperity ought to be much much bigger than that: it ought to be about 20 times richer.

The Bottom Line

Brad DeLong's commentary serves as a necessary reality check, exposing how a prestigious paper can persist by hiding behind complex statistics that crumble under scrutiny. The strongest part of his argument is the demonstration that the original study's results are not just weak, but mathematically incoherent, producing effects that are empirically falsified by the broader historical record. The biggest vulnerability, however, lies in the inertia of the field: even with this evidence, the narrative that "institutions are everything" remains dominant in policy circles. Readers should watch for whether the academic community will finally abandon these flawed instrumental variables or continue to defend a model that chimes thirteen.

The big picture from AJR (2001) remains intact and remarkably robust: Europeans were more likely to move to places that were relatively healthy, and when they moved in larger numbers, they imposed better institutions, which have tended to persist from the colonial period to today. But the big picture is a rhetorical position, not a statistical one.

Deep Dives

Explore these related deep dives:

  • Institutions, Institutional Change and Economic Performance Amazon · Better World Books by Douglass C. North

  • Instrumental variables

    Understanding the specific mechanics of instrumental variables is essential to grasp why the article argues that the 'first stage' in Acemoglu's study has collapsed and why the resulting estimates are statistically unreliable.

  • Cauchy distribution

    The article explicitly states the test statistic is 'near-Cauchy' with infinite variance, a technical detail that explains why standard significance tests fail and why the economic conclusions drawn from the data are mathematically unsound.

  • Colonial Origins of Comparative Development

    This specific dataset is the core empirical claim of the 'Colonial Origins' paper being debated, and the article details how the inclusion of conjectured rates and the omission of specific countries fundamentally alters the study's findings on development.

Sources

Wait! What?!?!: “Mr albouy reaches his conclusion by omitting half the data from the original…

The Economist published an unbylined, unsourced piece asserting Daron Acemoglu “counters that [David] Albouy reaches his conclusion by omitting half the data from the original sample.” That claim is simply wrong, and Albouy is right to be angry — no fact-check, no source, no context. Albouy’s actual point is narrow and correct: AJR have no real first stage. Once you correct for clustering, drop 36 conjectured mortality rates, and control for barracks-versus-campaign sources, the mortality–expropriation relationship collapses toward one-in-three significance. The second-stage test statistic isn’t a t-distribution; it’s near-Cauchy — infinite variance, no mean. The IV estimates are unreliable. Moreover, Acemoglu, Johnson, and Robinson ought to be very grateful to David If you take their IV results seriously, the effects implied are embarrassingly and implausibly large: a clock that chimes thirteen. Albouy provides an explanation for what is otherwise a very large implausibility in their story:.

One cannot know what to make of this paragraph in the London Economist:

Anonymous: The World’s Most Influential Economist Is Oddly Unconvincing <https://www-economist-com.libproxy.berkeley.edu/finance-and-economics/2026/08/17/the-worlds-most-influential-economist-is-oddly-unconvincing>: ‘David Albouy… showed that some countries were assigned mortality rates borrowed from other[s]…. Correct… and the [Acemoglu] paper’s estimates become unreliable…. Buchner… and colleagues reported that experts they surveyed were somewhat more likely to side with Mr Albouy. Mr Acemoglu… counters that Mr Albouy reaches his conclusion by omitting half the data from the original sample, including on important countries like America, Canada and Australia. It is this combination, along with some statistical choices, that introduces the unreliability, he says. He adds that if he were redoing the paper today, he would make “a number of changes, including in some of the estimation details”—though not to the mortality data…

To start with, the Economist’s lack of bylines makes hit pieces like this one on Daron Acemoglu unconvincing. The lack of sources does as well. Normally, you expect an unsourced “said” or “counters” to be something said to the reporter. But the story quotes a tweet from Noah Smith:

I’ve been yelling about Acemoglu for literally a decade…

And it did not contact Noah. It quotes a podcast segment from Larry Summers:

He leaves out entirely in that analysis the possibility that we will have more rapid scientific progress, more rapid social-scientific progress, or better decision-making because of artificial intelligence…

without stating the source as well.

Thus I have no idea what the context of the part of the story I take ...