← Back to Library

Can Amd break the cuda moat? Amd advancing AI 2026

In a landscape where Nvidia's dominance is often treated as an immutable law of physics, Dylan Patel of SemiAnalysis makes a startling pivot: he now believes AMD has a genuine shot at breaking the software moat that has protected the AI accelerator market for a decade. This isn't wishful thinking from a competitor; it is a data-driven re-evaluation from the industry's most rigorous bug hunter, who admits his own team spent months triaging broken code only to witness a dramatic cultural and technical turnaround. The piece is notable not for predicting a takeover, but for identifying the specific, non-obvious bottlenecks—internal testing capacity and supply chain friction—that will determine whether AMD's ambitious 2026 timeline survives contact with reality.

The Pivot from Skepticism to Strategy

Patel's narrative arc is perhaps the most compelling evidence of the shift he describes. He openly admits that just six months ago, his firm gave AMD a "0% chance of closing the gap with Nvidia in AI accelerators," citing software that was "broken" and progress that was "unexciting." Yet, the tone has shifted dramatically. "Based on our experience on AMD software stack this year, we update our view again from non-zero percentage chance to now a great chance of success," Patel writes. This is a rare admission of error from a top-tier analyst, and it carries weight because it is grounded in direct observation of leadership changes. He notes that CEO Lisa Su has "hopped on calls" with his team and implemented their suggestions, signaling a move away from "committee style leadership" toward an "agentic oriented engineering culture."

Can Amd break the cuda moat? Amd advancing AI 2026

The stakes of this shift are massive. Patel highlights that major players are already betting on this turnaround. "Anthropic has publicly announced that they will deploy 2GW of AMD's chips," he notes, a move that validates the hardware's potential despite historical software struggles. He points to Anthropic's Head of Compute, Tom Brown, who successfully used the "goal" command to bring up an internal inference stack on AMD hardware over a weekend. This anecdote serves as a powerful counter-narrative to the idea that switching away from Nvidia's ecosystem is impossible. "We believe that since AMD's compiler and most of their kernels are open sourced, that they are better positioned for the agentic age," Patel argues, suggesting that the open nature of the software stack allows for the rapid, iterative testing that AI agents require.

"The pie is growing rapidly for everyone, and Nvidia will continue to massively grow revenue. AMD poses potential competition to Nvidia on the software front."

However, Patel is careful not to frame this as a zero-sum game where Nvidia must fail for AMD to succeed. Instead, he suggests that the competition will force the incumbent to improve. He writes that "Jensen will need to cut bureaucracy and flatten the layers of different required internal stakeholder approvals required for even the simplest of tasks if he wants Nvidia to move faster and defend their lead." This reframing moves the conversation away from personality clashes and toward the institutional dynamics of large tech giants. It is a sobering reminder that in the AI race, the biggest threat to a monopoly is often the internal friction that comes with scale.

The Hardware Edge and the Software Trap

While the software narrative is the headline, Patel's deep dive into the silicon reveals why the hardware is finally catching up. The MI455X chip is described as a marvel of engineering, utilizing TSMC's 2nm process and a massive package size that allows for "12 stacks of HBM4 for a total of 432GB per package." This is a significant leap, especially when compared to competitors shipping only 8 stacks. Patel notes that this "silicon spam" delivers industry-leading theoretical performance, with the MI455X offering "20PF of FP8" compared to Nvidia's Rubin at 17.5 PF.

Yet, the hardware advantage is not a silver bullet. Patel points out that AMD is "forced to be aggressive on silicon to compensate for this deficiency" in microarchitecture design, specifically the lack of "3 bit LUT tensor cores" that Nvidia's Rubin possesses. This is a critical nuance: AMD is winning on raw silicon area and memory bandwidth, but Nvidia is still ahead in architectural efficiency. The memory bandwidth war is particularly fierce; while AMD boasts 23.3 TB/s, Patel reveals that Nvidia aggressively pushed pin speeds to 10.7Gbps to close the gap, resulting in a "barely above" 22TB/s for their own hardware. This cat-and-mouse game highlights how quickly the technological lead can evaporate when the competition is this close.

The economic model AMD is deploying is equally aggressive and unconventional. Patel details a "stock option based structure whereby AMD gives Meta and OpenAI close to a 105% equity rebate discount." He describes this as "clever financial engineering" that effectively makes the cost per million tokens "practically negative cost" for customers who hit certain equity targets. "AMD is practically giving away Helios racks and an 5% extra on top of that," he writes. This strategy is a direct attempt to overcome the inertia of the CUDA ecosystem by making the financial risk of switching to AMD virtually non-existent. Critics might note that such deep discounts could signal desperation or erode long-term profitability, but in a market defined by rapid scaling, the priority is clearly on securing deployment volume over immediate margins.

"The chief complaint from most AMD engineers internally is that there is a persistent lack of stable GPU clusters for internal software development teams and a lack of stable GPU clusters for automated testing CI."

The Hidden Bottleneck: Internal Capacity

Despite the hardware brilliance and the aggressive financial incentives, Patel identifies a single, existential risk that could derail AMD's entire 2026 roadmap: the lack of internal testing infrastructure. This is the piece's most critical insight. The argument is that software quality is not just a matter of code, but of the environment in which that code is tested. "Every time we meet with y'all... we always highlight that CI could be better," Patel recounts, noting that the planned timeline for parity with Nvidia's testing standards was missed due to "cluster issues."

The problem is exacerbated by the rise of "agentic coding." Patel explains that previously, a human engineer needed a few nodes to test code. Now, "each agent requires GPUs to test their code against and each human can have dozens of agents running at the same time." This exponential demand for compute is straining AMD's internal capacity, which remains "more than an order of magnitude less capacity than the stable long term clusters Nvidia has." The result is a vicious cycle where the lack of stable clusters prevents the software from improving, which in turn makes it harder to attract the talent needed to fix the software.

Patel is blunt about the consequences: "This is blocking AMD's rate of progress and it is holding AMD back from harnessing the potential upside of AI coding Agents." He points to specific regressions, such as the "vLLM gating automated test progress" which "massively regressed due to AMD cluster infra stability issues." The fact that leadership has been "pulling clusters away from AMD's internal vLLM team to deploy elsewhere" suggests a misalignment of priorities that could be fatal in a race where speed is everything. While AMD's Shanghai-based teams are doing "10x" work, they are being held back by a lack of resources. "We hope that AMD's leadership can re-prioritize giving their internal vLLM team stable clusters," Patel urges, framing this not as a technical issue but as a strategic imperative.

Bottom Line

Dylan Patel's analysis is a masterclass in separating the hype from the hard data, offering a nuanced view where AMD's hardware is undeniably world-class but its software success hinges on a single, fixable variable: internal testing capacity. The strongest part of his argument is the identification of the "cluster crunch" as the primary bottleneck, a detail often overlooked in favor of flashier specs. The biggest vulnerability, however, remains the timeline; if the internal infrastructure cannot be scaled fast enough to support the demands of agentic development, the "great chance of success" could evaporate before the 2026 deadline. For investors and industry watchers, the key metric to watch is not just chip shipments, but the stability of AMD's own internal CI systems.

"The lack of GPU nodes is a problem that has been getting even worse due to the rise of agentic coding."

The path forward is clear but difficult. If AMD can solve its internal capacity crisis, the combination of aggressive pricing, open-source software, and superior silicon could finally crack the Nvidia monopoly. If not, the hardware will remain a powerful but underutilized asset, a testament to engineering excellence that never reached its full potential.

Deep Dives

Explore these related deep dives:

  • High Bandwidth Memory

    The article cites Microsoft's 2023 departure due to 'unreliable Samsung HBM,' making this technical concept essential for understanding the supply chain fragility that previously doomed AMD's hardware.

Sources

Can Amd break the cuda moat? Amd advancing AI 2026

by Dylan Patel · SemiAnalysis · Read full article

When we published our first AMD software article, we gave AMD a 0% chance of closing the gap with Nvidia in AI accelerators. Software was broken, progress was unexciting, and we were the top bug submitter for many months with dozens of AMD engineers triaging our bug reports.

Six months later, in our AMD 2.0 article, we took the non-consensus position of upgrading from 0% chance to a much more meaningful chance at success. We published this opinion when sentiment towards AMD was sitting at rock bottom, and when most of the market thought we were were being far too optimistic towards AMD.

That view was based on our observations that AMD has the leadership that can create change rather than suffer from committee style leadership. Lisa quickly hopped on a calls with us and has since then implemented many of our suggestions. We saw an AMD that had finally recognized the importance of software and had a sense of urgency, moving in the right direction even if the destination was still far off. Since then, the signal has only gotten stronger.

Based on our experience on AMD software stack this year, we update our view again from non-zero percentage chance to now a great chance of success as long as AMD solves the two major risks we outline below. It is important to highlight that just because AMD gains market share, that doesn’t mean that Nvidia will do poorly. The pie is growing rapidly for everyone, and Nvidia will continue to massively grow revenue. AMD poses potential competition to Nvidia on the software front, and Jensen will need to cut bureaucracy and flatten the layers of different required internal stakeholder approvals required for even the simplest of tasks if he wants Nvidia to move faster and defend their lead.

Anthropic has publicly announced that they will deploy 2GW of AMD’s chips, Lisa Su and team have leaned into an agentic oriented engineering culture. Anthropic Head of Compute Tom Brown explained how he used Claude over the weekend with “/goal” to bring-up internal Claude inference stack on AMD hardware as a case in point. We believe that since AMD’s compiler and most of their kernels are open sourced, that they are better positioned for the agentic age besides one major risk that we will discuss later. Three months ago, our Accelerator model noted that Anthropic will be an AMD customer. ...