Google reportedly capped Meta's Gemini compute in March, delaying Meta safety tools
Google reportedly told Meta around March that it could not supply the full Gemini AI compute Meta had requested, delaying Meta's internal safety automation.
Google told Meta around March 2026 that it could not supply the full volume of Gemini artificial-intelligence compute Meta had requested, according to people familiar with the matter, delaying several of Meta’s internal AI projects.
The shortfall is one of the clearest signs yet that frontier AI capacity limits are disrupting even the largest operators. The cap reportedly hit Meta’s safety workflows hardest, the systems that use Gemini to automate harmful-content removal and scam detection. Meta had leaned on Gemini for that work because it outperformed the company’s own Llama models, making the restriction especially acute.
After the cap, Meta reportedly told staff to use AI tokens more efficiently and shifted workloads to Muse Spark, an internal model. Other Google Cloud clients were affected, but Meta was said to be hit hardest because its demand ran exceptionally high. The two companies compete directly in digital advertising, which makes the dependency unusually sensitive.
The constraint tracks what Google has said publicly about its own limits. Chief Executive Sundar Pichai told investors during first-quarter 2026 earnings that computing-power constraints held back even faster Google Cloud revenue growth; the unit posted $20 billion in revenue for the quarter. To add bridge capacity, Google signed a deal worth about $920 million a month with SpaceX for roughly 110,000 Nvidia graphics processing units. Google’s 2026 capital spending runs between $180 billion and $190 billion, while Meta has guided to $115 billion to $135 billion.
Neither Google nor Meta provided on-record comment, and the account rests on unnamed people familiar with the matter rather than any company disclosure. The figures on token efficiency and the scope of affected clients have not been independently verified.
The episode suggests that even hyperscalers with tens of billions in annual capex are rationing their best models internally, a dynamic that could push more workloads toward in-house systems like Muse Spark.
An entrepreneur with over a decade of experience in AI, Cloud, and HPC. He is currently a DevOps Architect and the founder of Data Phoenix, an influential media voice for the AI industry, with a strong focus on community building and open source.
More news

Google says AI pipeline helped Chrome fix 1,072 security bugs in June

Google, Constellation sign 20-year deal for 890 MW of nuclear power

Google to start limiting free Gemini users to Flash-Lite on October 9
