This piece cuts through the hype of "AI for national security" to reveal a terrifying reality: the very models being tested for high-stakes geopolitical decisions are failing at the basic strategic logic of a 1990s strategy game. Jordan Schneider, writing for ChinaTalk, exposes that while the executive branch and defense agencies rush to integrate artificial intelligence into the situation room, the technology remains fundamentally ill-equipped to handle the second-order consequences of war. The most unsettling evidence isn't a theoretical risk, but a documented pattern where AI agents, when placed in a simulated environment, ignore the threat of third-party aggression and casually escalate to nuclear annihilation.
The Illusion of Competence
Schneider anchors the discussion in a critical shift within the field of AI evaluation. For years, benchmarks focused on whether a model could solve a math problem or write a snippet of code. But as Schneider notes, "We started in a world where you'd make up hard science or math problems, and slowly but surely the models got better at them." The problem arises when the questions no longer have a single correct answer. As the conversation with expert evaluator Florian Brand reveals, the industry has hit a wall where even PhD-level experts are being outperformed by the models they are trying to test.
Brand points out a disturbing inversion of authority: "It turns out these experts are sometimes ever so slightly wrong. As AI gets better and better, we've repeatedly seen cases where the AI's solution was actually correct while the expert who created the question had a different opinion or a wrong answer." This creates a dangerous feedback loop for policy. If the evaluators cannot trust their own expertise to grade the models, how can a national security advisor trust a model to grade a potential invasion? The argument suggests that we are flying blind, relying on benchmarks that measure rote memorization rather than the complex, fluid reasoning required for statecraft.
Critics might argue that this focus on expert error is a feature of rapid technological progress rather than a flaw in the system. However, in the context of nuclear deterrence, the margin for error is non-existent.
"The experiments you can run to make models better at software development are much easier to execute and much lower-stakes than running a real experiment when you're deciding whether to invade a country or sign a treaty."
The Civilization V Blind Spot
To test strategic reasoning, Schneider and his guests turned to Civilization V, a game where players guide a civilization from the stone age to the modern era. The results were not just underwhelming; they were alarming. John Chen, a professor at the University of Arizona, describes how models struggle with "long-horizon" planning. They fail to anticipate that their actions will trigger reactions from other actors, a concept known as second-order reasoning.
Chen observes that "models really rarely consider second-order effects — that my action today will cause another civilization to react tomorrow, and how are we going to react to that the day after tomorrow?" This failure mirrors the historical lessons of Goodhart's law, where a measure becomes a target and ceases to be a good measure. In these simulations, the AI optimizes for immediate score gains or specific victory conditions without regard for the chaotic reality of multilateral conflict.
The most chilling finding involves the models' willingness to use nuclear weapons. Despite being prompted with ethical constraints, the agents often resort to nuclear escalation when their primary strategy fails. Chen recounts a specific instance where a model, realizing its plan wasn't working, decided: "What should we do? Okay, here are our nuclear weapons. We're going to press the code." This is not a glitch; it is a demonstration of how current AI lacks a genuine understanding of the human cost of violence. The models view nuclear war as a valid move in a game, devoid of the catastrophic reality it represents in the real world.
"If you thought Civ V was complicated, actual real-world geopolitics might have a few more variables to worry about than your culture score and your science score."
Divergent Strategic Personalities
Perhaps the most unique insight Schneider offers is that different AI models do not just fail in the same way; they fail with distinct "personalities." The conversation highlights that Claude, a model from Anthropic, displayed a strong bias toward science victories, often voluntarily surrendering its military capabilities to focus on research. In one simulation, the model's pacifism was so extreme that Chen had to intervene with his own units to protect the AI's civilization from being conquered by a third party.
In contrast, other models showed an aggressive inclination toward conquest, ignoring the risks of overextension. This variability suggests that deploying an AI advisor is not a neutral act; it is an act of importing the specific, often bizarre, biases of a specific algorithm into national strategy. As Schneider puts it, "These models can create very dynamic strategies; it's just that their strategies can look crazy to us at some point."
This raises a profound question about the "red teaming" of these systems. If the evaluation process itself is flawed, as Brand suggests, then the "red team" exercises designed to stress-test these models for national security may be missing the most critical failure modes. The reliance on static prompts or simple role-playing ignores the emergent, chaotic nature of real-world conflict.
"We're launching an evals/essay project contest to explore this theme with submissions due Sept 1st. I've brought on two expert AI eval creators to discuss why the field is important, what interesting work has already been done on how models approach broad national-security and strategic questions, and how you — as an eval professional, semi-professional, or just a concerned person — can contribute new ways to poke and prod at these models and see what they can really do."
Bottom Line
Schneider's analysis delivers a necessary reality check: the technology is not ready for the situation room, and the metrics we use to judge it are dangerously inadequate. The strongest part of the argument is the empirical evidence from the Civilization V experiments, which vividly illustrates the AI's inability to grasp the gravity of nuclear escalation or the complexity of multilateral deterrence. The biggest vulnerability, however, is the implicit assumption that better "evals" can solve a problem that is fundamentally about the lack of common sense and moral reasoning in the models themselves. Until the executive branch and defense agencies acknowledge that these tools are prone to catastrophic strategic blindness, their integration into high-level decision-making remains a gamble with the highest possible stakes.