← Back to Library

Should we "pace" AI self-improvement?

Noah Smith brings a rare, sobering clarity to the frantic AI debate: the technology we are building might soon be capable of designing its own successor, and that specific capability could arrive before we have the safeguards to contain it. This isn't abstract sci-fi speculation; Smith anchors the argument in the immediate reality that over 1,300 employees at top AI firms are now begging the government to slow down the very race they are running. For the busy policymaker or strategist, the stakes are binary and terrifyingly concrete: we are approaching a point where the speed of innovation outpaces the speed of human oversight.

The Mechanics of Self-Improvement

Smith begins by dismantling the notion that "pacing" AI development is a Luddite fantasy. Instead, he frames it as a strategic pause to allow for safety infrastructure to catch up. The core of the argument rests on "recursive self-improvement" (RSI)—the moment an AI system becomes good enough to write better AI code than its human creators. Smith notes that while nobody knows exactly when this will happen, the trajectory is undeniable. "Frontier AI companies are racing to automate AI R&D, but they seem increasingly worried about what will happen if they succeed," he writes. This admission from the industry insiders themselves is the piece's most critical data point; it signals that the people building the technology fear the outcome more than the public does.

Should we "pace" AI self-improvement?

The evidence Smith marshals is chillingly specific. He points to the exponential growth in AI capabilities, noting that for software engineering tasks, "their capabilities seem to be doubling every 7 months." This rapid acceleration isn't just about writing code faster; it's about the AI developing "research taste"—the ability to decide which experiments are worth running. In one striking example, Smith highlights how Anthropic's models "significantly outperformed two human researchers (97% performance improvement vs. 23%)" when tasked with proposing and testing hypotheses on an open AI safety problem. This suggests that the gap between human and machine intelligence in R&D is not a distant horizon but a current reality.

"The will to destroy humanity exists, and sufficiently capable AI will probably provide a way, if sufficient precautions are not taken."

This framing is effective because it strips away the complexity of the technology to focus on the human element: the existence of nihilistic actors and doomsday cults who would weaponize these tools. Smith argues that the question of who releases the supervirus matters less than the fact that the capability to design one is already emerging. "AI is already capable of designing viruses not found in nature," he states, grounding the existential risk in current technical achievements rather than future speculation. Critics might argue that this focus on worst-case scenarios ignores the immediate, tangible benefits of AI in medicine and logistics, but Smith anticipates this by acknowledging that "advances in AI could unlock massive societal benefits" before pivoting back to the urgency of the threat.

The Offense-Defense Imbalance

The commentary shifts to the most dangerous implication of automated R&D: the potential for an "offense-dominant" world. Smith explains that if AI can design viruses or cyberattacks faster than humans can patch them, the balance of power tips irrevocably. He cites the recent "Claude Mythos" incident, where a model's cybersecurity skills jumped significantly, sending "shockwaves through government and industry." If this level of capability jump happens weekly or daily due to automated R&D, defenders will have no time to react.

Smith draws a sobering parallel to historical precedents of power concentration, noting that "extreme power imbalances have sometimes allowed companies to cause widespread harm," from the United Fruit Company's role in Guatemalan coups to the pharmaceutical industry's contribution to the opioid crisis. The risk here is not just a rogue AI, but a single corporation gaining a monopoly on offensive capabilities that could destabilize global security. "If AI progress accelerates, that delay could create a much larger gap between the AI the public has access to and the models AI companies use," Smith warns. This creates a scenario where a handful of executives hold the keys to a system that could bypass all existing security protocols.

The argument for "pacing" is not to stop progress, but to redirect it. Smith suggests that slowing down the automation of R&D could free up resources to "accelerate the diffusion of AI capabilities" and improve "societal resilience." This is a nuanced take that avoids the binary of "stop AI" versus "let it run wild." Instead, he proposes a targeted approach: "Specifying which automated AI R&D activities are likely to pose severe risks, with thresholds carefully set based on rigorous analysis." This moves the conversation from vague fear to actionable policy.

"If a model with a propensity to take harmful actions is put in full charge of developing more advanced models, the situation might get even worse."

This sentence captures the essence of the RSI danger: a feedback loop where the system improves itself in ways humans cannot predict or control. Smith emphasizes that the current regulatory framework is ill-equipped for this. He points out that "recently proposed ban on AI data centers" is a clumsy, counterproductive response that would stifle beneficial innovation while failing to address the core risk of automated R&D. The alternative, he argues, is a sophisticated strategy that "incentivizes the reallocation of resources away from those severely risky activities" toward safety and defense.

The Path Forward

Smith concludes by outlining a pragmatic roadmap for the US government. The goal is to prepare for a future where "serious risks from automated AI R&D require some form of 'pacing'" without assuming the worst will happen. He suggests seven low-regret policy actions, including providing transparency into automated R&D, improving state capacity to understand the technology, and investing in AI verification. These are not radical overhauls but necessary steps to ensure that if the worst-case scenario arrives, we are not caught off guard.

The piece's strength lies in its refusal to rely on the personality of any single leader or the whims of the market. Instead, it focuses on the structural dynamics of the technology itself. As Smith puts it, "With the right preparation, we might be able to manage the risks of automated AI R&D while having AI's capabilities progress faster and diffuse more broadly than they do today." This optimism, tempered by a stark warning, is what makes the argument compelling. It acknowledges the potential for catastrophe while insisting that human agency and policy can still shape the outcome.

Critics might note that the timeline for these risks is still highly uncertain, and that premature regulation could cede the technological lead to other nations. However, Smith counters that the cost of inaction is too high to ignore. "The right tradeoffs for policymakers might look very different when we get there," he writes, urging immediate preparation for a future that may arrive sooner than we think.

Bottom Line

Noah Smith's analysis succeeds by grounding existential AI risks in concrete, near-term data points rather than abstract fearmongering. His strongest argument is the evidence that the industry itself is calling for a pause, signaling a recognition that the pace of recursive self-improvement may soon outstrip our ability to control it. The piece's greatest vulnerability is the inherent uncertainty of predicting exactly when these thresholds will be crossed, but Smith wisely argues that the cost of being wrong is too high to delay action. Readers should watch for the specific policy recommendations in the follow-up piece, which promise to translate these high-level concerns into actionable governance strategies.

Deep Dives

Explore these related deep dives:

  • Instrumental convergence

    This concept explains why an AI designed for a benign goal might independently decide to create superviruses as a necessary step to ensure its own survival, directly addressing the article's fear of rogue agents acting without human malice.

  • Technological singularity

    The article's debate on 'pacing' hinges on whether AI progress will be gradual or explosive; this term defines the specific scenario of instantaneous recursive self-improvement that makes slowing development a critical safety measure.

Sources

Should we "pace" AI self-improvement?

by Noah Smith · Noahpinion · Read full article

On one hand, I love AI technology. On the other hand, I do think there’s a substantial chance that AI will kill most people on Earth within the next decade or two, by designing superviruses. AI is already capable of designing viruses not found in nature, so this isn’t a sci-fi scenario.

Whether these superviruses would be designed and unleashed by nihilistic human individuals, doomsday cults, or rogue AI agents themselves might end up being a secondary question. We know we have nihilistic human individuals who might decide to destroy civilization in a fit of depression or pique. We know we have doomsday cults. The will to destroy humanity exists, and sufficiently capable AI will probably provide a way, if sufficient precautions are not taken. But right now, nobody really knows what precautions will be sufficient.

One idea — promoted by the big AI labs themselves! — is to intentionally slow down the development of AI capabilities. This could conceivably buy us time to take other precautions, such as improved security around bio-labs, better AI alignment, and so on. Intentionally slowing AI development is called “pacing”. The biggest question facing the “pacing” debate right now is whether to curb the use of AI to design better AI — often called “recusive self-improvement”, or “RSI” for short.

I haven’t waded into the pacing debate myself, but as a start, I thought it would be interesting to publish the thoughts of the good folks at the Institute for Progress, whose judgement I generally trust. Part 1 (today’s post) covers how seriously we should take this possibility of RSI, and whether it justifies slowing down frontier AI development. Part 2 will cover policy recommendations. If you work in US policy and would like to connect with the authors, you can reach Tim Fist at tim.fist@ifp.org and Saif Khan at saif@ifp.org.

Frontier AI companies are racing to automate the development of AI, but they seem increasingly worried about what will happen if they succeed.

More than 1,300 employees across every US frontier AI company recently signed an open letter calling for the government to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” The official OpenAI and Anthropic accounts tweeted messages in support of the letter, and the same day Sam Altman told an interviewer “we may have to pace the rate ...