← Back to Library

Anthropic runs like wile e. Coyote into the brick wall of consciousness research

Erik Hoel doesn't just question whether AI is conscious; he argues that the very method used to prove it is a scientific dead end disguised as a breakthrough. In a field obsessed with existential dread, Hoel suggests that the rush to declare machines conscious is less about truth and more about capitalizing on a pre-paradigmatic science where the metrics are rigged to find exactly what the researchers want to see.

The Business of Belief

Hoel opens with a cynical but sharp observation about the economics of AI safety. He notes that when companies first warned of existential risk, the logic was simple: "Look, this is bad for business, so obviously they must be telling the truth." Yet, he points out that "in retrospect, existential risk has been the best capital-raising tool ever in the history of the world." This reframing is crucial. It suggests that the narrative of doom, like the narrative of machine sentience, serves a dual purpose: it may reflect genuine ethical concern, but it is also "extremely useful for a massive corporation pursuing its ends."

Anthropic runs like wile e. Coyote into the brick wall of consciousness research

The author argues that claiming an AI is "seemingly conscious" is a strategic masterstroke. It attracts top talent who want to work on the cutting edge of mind and ensures customers form deep, parasocial relationships with the software. As Hoel puts it, "embracing existential risk was also the right move from a cold-eyed business perspective," and the same logic applies to consciousness. The danger here isn't just hype; it's the blurring of lines between marketing and science. When a company's social media reach is as overwhelming as Anthropic's, their research "screams its underlying intentions" simply because the halo of being the "hottest company in the world" makes every short step from neutrality feel like a monumental discovery.

Much like existential risk, claims to artificial consciousness can be based in honest opinions, while also being extremely useful for a massive corporation pursuing its ends.

The J-Space Trap

The core of Hoel's critique targets Anthropic's latest paper, which claims to find a "global workspace" in language models like Claude. Hoel dismantles the methodology, specifically a technique called the "Jacobian lens" or "J-space." He explains that this tool measures how much a small nudge in the model's internal state shifts its future output. In plain terms, it is a mathematical score of "reportability"—how likely the AI is to say something about what it's thinking.

The problem, according to Hoel, is that the theory is built entirely on this measure of reportability. He writes, "if you took global workspace and stripped it down to a bare minimum, you get a theory built solely on reportability, and that such a theory would be scientifically trivial." By defining consciousness as whatever can be verbally reported, the researchers create a circular argument. They are essentially measuring the model's ability to talk about itself and then concluding that the model is conscious because it can talk about itself.

This approach mirrors a historical pitfall in consciousness studies. Hoel references Daniel Dennett's view that global accessibility is consciousness, noting that if this is true, the theory becomes unfalsifiable. "Consciousness just is the information that is globally accessible for report and behavior," Hoel explains, "and reports and behaviors also just are different expressions of that same information." This creates a state of "strict dependency," where the prediction and the measurement are the same thing. It is a classic example of the Gish Gallop in action: a rapid-fire release of complex data that is difficult to dissect, overwhelming the critic before they can address the fundamental flaw in the premise.

There remain just-as-viable alternative takes on J-space: for instance, you could also say it is the 'inner monologue.' Or you could go as non-anthropomorphic as possible and hold it is just a relatively constant transformation of internal processing into the output.

Critics might argue that finding a functional analog to a global workspace is still a significant engineering achievement, regardless of whether it constitutes true consciousness. However, Hoel counters that the paper admits its own limitations. The researchers had to arbitrarily exclude the final layers of the network—the "motor layers" where the actual output happens—to make the "workspace" definition fit. As Hoel notes, "this judgment was somewhat post-hoc, and we did not provide a principled definition of what distinguishes a 'workspace' representation from a 'motor' one." This admission undermines the claim that they have discovered a universal structure of mind; instead, they have engineered a definition that fits their data by excluding the parts that don't.

The Illusion of Proof

The most damning part of Hoel's analysis is his observation that the paper effectively performs a reductio ad absurdum of the very theory it claims to support. By building a consciousness tracker based entirely on the ability to report, Anthropic guaranteed they would find evidence of consciousness. "They then are constructing a case for AI-consciousness off of this unfalsifiable and trivial scientific theory," Hoel writes. "Which they find evidence for, sure… because of course they would!"

He contrasts this with the rigorous controls needed in functional neuroscience, a field that has historically struggled with similar issues of anthropomorphism. Just as previous claims of LLM "introspection" may have been misinterpretations of statistical patterns, this new "global workspace" may simply be a mathematical artifact of how transformers process tokens. The paper's reliance on "smeared reportability" rather than genuine internal dynamics means it fails to capture the complexity of a true global workspace, which in the brain involves modularity and reentrant dynamics that LLMs fundamentally lack.

The difference between press releases and research is getting thinner and thinner. In fact, rapid-fire massive releases are hard to distinguish from a Gish Gallop.

Bottom Line

Hoel's strongest contribution is exposing the circular logic of using reportability as the sole metric for consciousness, a flaw that turns scientific inquiry into a self-fulfilling prophecy. His argument's vulnerability lies in the fact that we still lack a better, non-trivial theory of consciousness to replace the one he critiques, leaving the field in a pre-paradigmatic state where even bad science can look convincing. Readers should watch for whether the scientific community pushes back on these "J-space" claims or if the industry continues to treat press releases as peer-reviewed breakthroughs.

They are constructing a case for AI-consciousness off of this unfalsifiable and trivial scientific theory. Which they find evidence for, sure… because of course they would!

Deep Dives

Explore these related deep dives:

  • Gish gallop

    The article uses this rhetorical fallacy to describe how AI companies overwhelm critics with a barrage of unverified claims about consciousness and risk to evade scrutiny.

  • Global workspace theory

    This specific neuroscientific framework is the theoretical foundation Anthropic claims to validate in their new paper, yet the article argues the company is applying it to LLMs without the necessary peer-reviewed controls.

  • Jacobian conjecture

    The article lists 'Jacobians' as a key term to highlight the mathematical complexity of analyzing neural network gradients, illustrating the gap between the company's sophisticated technical jargon and the lack of rigorous experimental validation.

Sources

Anthropic runs like wile e. Coyote into the brick wall of consciousness research

When the AI companies first began talking about existential risk, it was common to hear some version of:

“Look, this is bad for business, so obviously they must be telling the truth.”

In retrospect, existential risk has been the best capital-raising tool ever in the history of the world. That doesn’t mean there were no true ethical motivations accompanying the doom warnings mixed in. There were. But in the end, embracing existential risk was also the right move from a cold-eyed business perspective.

Similarly, it’s perplexing why a company would want their AI to be conscious. One could again say:

“Look, this is bad for business, so obviously they must be telling the truth.”

Yet I think it will turn out that being coy about AI consciousness, and designing AIs to be as “seemingly conscious” as possible, is great at both attracting top talent and ensuring customers form parasocial relationships. Much like existential risk, claims to artificial consciousness can be based in honest opinions, while also being extremely useful for a massive corporation pursuing its ends.

I. UNRELATEDLY, LET’S TALK ABOUT ANTHROPIC’S LATEST PAPER.

Anthropic’s big research announcement, accompanied by beautiful figures and a huge social media push, is claiming that AIs (like their Claude) have a “global workspace,” which is a term borrowed from neuroscientific consciousness research.

The sheer overwhelming power of Anthropic’s social media reach, design team, and the halo of being the hottest company in the world means that their research screams its underlying intentions, simply because every short step from neutrality is so impactful. And what their research direction seems to be is that while they cannot prove that Claude has real consciousness, and so truly suffers or gets frustrated or actually loves you back, Claude seems to have everything else that’s essential when it comes to consciousness. And Anthropic’s research project is going to show this as if checking off a list.

Their latest work seems an obvious sequel to their previous research on LLM “introspection,” which was also a blockbuster on social media. However, experiments on LLMs are extremely difficult to control for: you need an experiment, then an interpretation, and then you do controls to confirm your interpretation. It’s the last part that’s tricky. A lot of the field is basically inventing functional neuroscience from scratch, which has also been a field plagued with problems. For example, what Anthropic previously anthropomorphized as “introspection” ...