← Back to Library

Meta’s infrastructure team needs a culture reset

Most industry analysis treats Meta's hardware spending as a straightforward arms race for artificial intelligence dominance, but Dylan Patel argues the real story is a self-inflicted wound driven by internal dysfunction. In a scathing assessment of the company's infrastructure division, Patel contends that political maneuvering and a toxic performance culture are costing the tech giant billions, turning it into a cautionary tale of over-engineering rather than a blueprint for efficiency. This is not just about bad code; it is about a system where middle managers protect their turf by building unnecessary complexity, leaving the company's world-class researchers stranded with inferior tools.

The Culture of Short-Termism

Patel's central thesis is that Meta's infrastructure teams have become bloated and politically charged, prioritizing the appearance of activity over the delivery of usable technology. He writes, "Middle managers will do everything to justify their proposals to protect their positions within Meta, which has become an extremely political organization." This observation strikes a chord because it explains why a company with such immense resources keeps making baffling technical choices. The author points to the company's six-month performance review cycle, which cuts the bottom 10% to 15% of employees every round, as a primary driver of this behavior. "The result is an organization of employees that optimize for short-term wins rather than long-term strategies," Patel notes, describing a phenomenon where managers pursue "window washing" projects that look good quickly but are abandoned before they can deliver real value.

Meta’s infrastructure team needs a culture reset

This dynamic creates a environment where risk aversion reigns, and few feel safe challenging leadership decisions. The author draws a parallel to the concept of Goodhart's law—the idea that when a measure becomes a target, it ceases to be a good measure—suggesting that Meta's obsession with specific metrics has blinded the organization to broader needs. Critics might argue that high-pressure environments are necessary for rapid innovation in the AI race, but Patel's evidence suggests the opposite: the pressure is causing the organization to fracture and waste capital.

Middle managers will do everything to justify their proposals to protect their positions within Meta, which has become an extremely political organization.

The Rivos Acquisition: A Case Study in Dysfunction

The most damning evidence Patel provides is the $2.5 billion acquisition of chip startup Rivos. He frames this not as a strategic masterstroke, but as a chaotic transaction driven by a desire to own intellectual property rather than a clear technical roadmap. "Few inside Meta's chip division have a full understanding of why the company bought Rivos in the first place," he writes, noting that the deal was likely pushed by leadership simply because "Meta had the money" and the custom silicon space was heating up. The result was a disaster where the startup's founders insisted on an all-or-nothing deal, leading Meta to acquire the entire company only to immediately cut the parts they didn't want.

Patel describes a culture where existing managers treated the acquired engineers as a "pool of free headcount," pulling them apart to build their own empires until "little of it remained intact." This mirrors the historical trajectory of many failed tech integrations, where the lack of a unified vision leads to the rapid erosion of talent. The author highlights that despite the massive cost, the original technology goals evaporated; the chip project meant to use Rivos's GPU architecture was cancelled, and the team is now working on a new project, Phoebe, with uncertain prospects. "Unlike more mature hardware organizations like Apple, Meta's hardware roadmaps can change just as quickly as its software roadmaps," Patel observes, noting that employees often find themselves working on projects that become pointless within six months. This is a stark contrast to the stability required for complex hardware development, where years of planning are the norm.

Hardware Missteps: Grand Teton and Ariel

The cultural rot extends beyond personnel to the physical hardware itself. Patel dissects Meta's server designs, specifically the "Grand Teton" and "Ariel" configurations, as examples of over-optimization driven by siloed teams with conflicting goals. The Grand Teton server, for instance, added a complex switch tray to house extra storage, a move Patel argues was unnecessary. "In production, the storage wasn't utilized nearly as much as anticipated by the model teams, which is why this design ended up being cancelled," he explains. The decision was driven by a desire to avoid reliance on Nvidia networking gear, yet the software stack still depended on Nvidia's protocols, rendering the hardware changes futile.

The situation worsened with the "Ariel" server design for the Blackwell generation. Meta opted for a configuration with one GPU and one CPU per board, rather than the industry standard of two GPUs, to better suit its recommendation systems. "Meta was the only customer of this Ariel SKU," Patel writes, pointing out that this choice resulted in a total cost of ownership that was 14% higher than the standard configuration. The author argues that this was a classic case of one division optimizing for its own needs at the expense of the company's broader AI ambitions. "This decision cost Meta billions of dollars," he states bluntly, noting that the LLM teams were left with an inferior system that increased costs while reducing performance for generative AI workloads.

This analysis suggests a fundamental disconnect between the infrastructure team and the research teams they are supposed to serve. Patel notes that the infrastructure team is "burdened by far too many disparate groups that are over-optimizing for certain metrics as opposed to delivering usable technology for the company as a whole." The irony, he points out, is that the company's attempt to reduce reliance on Nvidia actually increased their dependence on Broadcom switches, while failing to solve the stability issues they hoped to address. A counterargument might be that Meta's unique workloads require unique hardware, but Patel's data on the higher costs and lower utilization rates suggests these were not necessary trade-offs.

This decision cost Meta billions of dollars.

Bottom Line

Dylan Patel's analysis offers a piercing look at how internal politics can derail even the most well-funded technological ambitions, proving that culture is as critical as code. The strongest part of his argument is the detailed breakdown of how specific cultural incentives—like the short-term performance reviews—directly led to catastrophic hardware decisions like the Rivos acquisition and the Ariel server. The biggest vulnerability in the piece is the lack of a clear path forward for Meta to implement the "cultural reset" he demands, given the entrenched nature of the political dynamics he describes. For investors and industry watchers, the takeaway is clear: until Meta addresses the disconnect between its infrastructure managers and its AI researchers, its massive capital expenditures may continue to yield diminishing returns.

Deep Dives

Explore these related deep dives:

  • Goodhart's law

    The article describes how Meta's infrastructure teams over-optimize for specific metrics to satisfy middle managers, a classic manifestation of this economic principle where a measure becomes a target and ceases to be a good measure.

Sources

Meta’s infrastructure team needs a culture reset

by Dylan Patel · SemiAnalysis · Read full article

In our recent newsletter piece about Meta Superintelligence, we expressed reasons to be optimistic on Meta AI. MSL now has many of the right ingredients to catch up with Anthropic and OpenAI to return to the frontier. However, we also briefly alluded to cultural issues plaguing Meta’s infrastructure teams. This article will dive into how these cultural issues have manifested into expensive missteps, whether it be with acquisitions like Rivos or strange choices on hardware architecture.

We believe that Meta Infrastructure needs a cultural reset to better serve the Meta AI organization, especially the world class researchers who are at MSL. This is even more important as Meta embarks on the path of selling its compute to outside customers, not just serving captive internal users.

Meta Infrastructure has become bloated, with middle managers expending resources on over-engineered technology solutions that lose sight of broader organizational needs. The company appears burdened by far too many disparate groups that are over-optimizing for certain metrics as opposed to delivering usable technology for the company as a whole. Middle managers will do everything to justify their proposals to protect their positions within Meta, which has become an extremely political organization.

One big issue is Meta’s six-month performance review cycle, in which the bottom 10% to 15% are cut every review round. The result is an organization of employees that optimize for short-term wins rather than long-term strategies. Some managers push for highly visible projects that can be delivered quickly, a practice known as “window washing” and then promptly pivot or abandon them. Few openly challenge leadership, which leads to bad decisions going uncorrected. The whole system discourages long-term thinking and leads to risk-adverse behavior.

Within Meta Infrastructure, supply chain teams also have little say over engineering teams. The result is technology decisions driven by political motivations rather than thoughtful software/hardware co-design for the broader company.

Frequent pivots are also common. And because Meta has a reputation for throwing money at problems and executing at high speed, these U-turns end up becoming more costly versus other companies that take a more disciplined or conservative approach. Suppliers also lose faith when given design wins are later cancelled. This has lead to less supply chain prioritization on new designs. Some suppliers favor focusing on Amazon or Google designs due to Meta’s frequent reshuffling.

A lot of Meta’s issues come from a lack of financial discipline, with managers ...