NewDust announces Series B to fuel next chapter of growth

The Compute Availability Trap: Why Multi-Provider AI Is No Longer Optional

Thibault MartinThibault Martin
-June 9, 2026
Why Multi-Provider AI Is No Longer Optional
"There aren't enough chips; not enough memory, motherboards, hard drives; not enough helium for semiconductors; not enough electrons. The entire semiconductor supply chain is under pressure from AI." — Arthur Mensch, CEO of Mistral AI, testimony before France's Assemblée Nationale, May 12, 2026 [8]
Mensch is describing a resource constraint. And when a critical resource becomes scarce, how you source it starts to matter as much as how you use it.
A few months ago, we wrote about the risks of building on a single AI provider [0], focused on preserving access to the frontier as it moves. We mentioned outages, pricing changes, and capacity constraints in passing. Since then, every one of them has materialized — with names and dates attached.
Anthropic's API dropped below 99% uptime. OpenAI killed a product line overnight to redirect compute, taking down everything built on it. Usage caps appeared across Anthropic, Google, and GitHub Copilot. The underlying reason became clearer: AI providers are running out of physical capacity to serve demand, and they are making hard allocation decisions as a result.
Intelligence is becoming a utility — like electricity. And the lesson from every other utility is the same: you don't source it from a single supplier. You build a portfolio. Some sources are cheaper, some more reliable, some sovereign. When one goes down, the others absorb the load. The organizations that treat their AI supply this way will be the ones still running when the next capacity crunch hits.

TLDR

  • Compute infrastructure is now the primary constraint on AI availability. Demand is outpacing physical capacity, and providers are rationing access.
  • The evidence is already concrete: Anthropic's API uptime dropped below 99% in Q1 2026. OpenAI shut down Sora overnight to redirect compute. NERC issued a Level 3 alert identifying data center load as a systemic grid risk.
  • Single-provider platforms absorb every outage, every price hike, and every product cancellation in full. Multi-provider platforms have options.
  • Organizations choosing AI platforms today are making bets that take months to unwind. The time to build optionality is before you need it.

The Bottleneck Moved

AI model quality has improved fast and keeps improving. But demand is growing faster than the physical infrastructure to serve it. OpenAI's APIs processed 6 billion tokens per minute in October 2025. By April 2026, that figure had reached 15 billion — a 2.5x increase in five months [4]. Anthropic reported annualized revenue growth of roughly 80x in Q1 2026, far outpacing the compute capacity they had planned for [5].
This shows up in concrete ways:
  • Usage caps. Anthropic introduced weekly usage caps on Claude. Google cut Gemini API free-tier quotas significantly [6]. GitHub Copilot introduced weekly limits, with users reporting that heavy usage exhausted allowances within days [7].
  • Product cancellations. OpenAI shut down Sora on March 24, 2026. It was losing roughly $1 million per day in compute costs against $2.1 million in total lifetime revenue [2]. The $1 billion Disney deal for Sora integrations collapsed. Any product built on that API ceased to exist.
  • Model router rollbacks. OpenAI rolled back its automatic model router for free users because routing them to reasoning models was too expensive at scale.
GPU procurement timelines run 36 to 52 weeks from order to deployment. Grid interconnection queues in the US can stretch eight years or more. No amount of funding or engineering velocity can sprint past a 52-week lead time or an eight-year grid queue. 

What Happens When You Bet on One Provider

The Anthropic-SpaceX deal announced on May 6, 2026 is the clearest signal yet [1]. Anthropic committed $1.25 billion per month — $15 billion per year — for access to SpaceX's Colossus data center in Memphis. The deal runs through May 2029.
Two things matter here. First, Anthropic CEO Dario Amodei had publicly described Elon Musk's behavior as disqualifying him from positions of public trust. SpaceX is Musk's company. Anthropic signed the deal anyway. When your compute options are constrained enough, ideology yields to necessity. Second, this happened because Claude demand grew far beyond what Anthropic had capacity to serve. The pressure was already visible: Anthropic's Claude API recorded 98.95% uptime over the 90 days ending April 2026 — roughly 3.5 days of downtime in a single quarter [9].
For a team shipping on a monthly sprint cadence, 3.5 days of downtime in a quarter is significant. For a team that has built a customer-facing product on Claude, it is a support crisis.
Even Anthropic itself is diversifying its compute supply — signing deals with multiple data center providers, securing capacity across different facilities. Providers understand portfolio logic for their own infrastructure. The question is whether you apply the same logic to how you consume their output.
If you are on a single provider and that provider hits capacity limits, you absorb every outage with no fallback. Your power users hit weekly caps right when they need the platform most. If the provider kills a product to redirect compute, everything you built on it disappears immediately. None of these require a catastrophic failure. They accumulate quietly, degrading reliability and user trust over months.

Your Pricing Leverage Disappears

Providers signing billion-dollar infrastructure deals need to pay for them. When your sole AI provider raises token prices to fund a $15 billion per year compute arrangement, you absorb that cost increase in full. There is no negotiation point and no alternative.
Multi-provider access changes this equation. When you run meaningful volume across Anthropic, OpenAI, Google, Mistral and a wide range of open-source models, each provider knows your usage can shift. That creates real pricing pressure. When one raises prices, you can route more traffic to the others while the situation settles.
A portfolio approach also unlocks cost optimization that a single-provider setup simply cannot offer. Route simple tasks to fast, cheap models. Reserve expensive reasoning models for the work that actually requires them. Match the cost of intelligence to the value of the task. That kind of allocation — the same logic enterprises already apply to cloud spend, energy procurement, and every other utility — is only available when you have more than one model in production.
At Dust, when Claude pricing changed, having GPT-5 and Gemini as active alternatives in production meant the platform could adapt without passing a cost spike to users. A single-provider platform takes a 30% price increase as a 30% hit to AI operating costs. A multi-provider platform has room to move. 

Cheaper Tokens Do Not Solve This

A reasonable objection: inference costs have dropped roughly 10x per year for comparable capability. If models keep getting cheaper, does compute scarcity actually matter?
It does, because cheaper tokens generate more usage. This is Jevons Paradox applied to AI: as cost per token falls, the number of tokens consumed rises — often faster than the cost reduction. Token demand at OpenAI grew 2.5x in five months even as per-token prices stayed flat or declined [4].
If falling inference costs were resolving capacity constraints, providers would not be signing trillion-dollar infrastructure deals. They would not be rationing access. The efficiency improvements are real, but demand is outpacing them. NERC issued its Level 3 alert — the highest urgency level — in early May 2026, identifying data center load growth as a systemic risk to grid stability [3]. Grid interconnection queues already stretch years. The physical infrastructure gap is durable, and it is widening on the timescale that matters for product decisions being made today.

The Conflict of Interest Gets Sharper Under Scarcity

When compute is plentiful, a lab that builds both models and a platform can serve all customers without friction. When compute is scarce, a choice emerges: whose requests get priority?
OpenAI's decision to redirect Sora's compute budget to other products is the clearest example. The choice was rational from OpenAI's perspective. It was catastrophic for anyone who had built on Sora.
Google, Anthropic, and OpenAI are all building their own agent platforms and consumer products. These products compete with third-party applications for the same compute capacity. When that capacity is constrained, providers make allocation decisions. Your API calls sit somewhere in a priority stack, and you do not get to see how it is ordered.
At Dust, this conflict does not exist. Dust has no model to protect and no compute to allocate. The only optimization is toward what works best for users. When Anthropic had outages in 2025, Dust customers continued working because the platform routed to OpenAI or Google automatically [10]. During a recent demo, the default model went down mid-conversation. The person presenting opened the agent builder, switched to OpenAI, and kept going. That moment — undramatic, seamless — is the whole point.
Multi-provider routing sounds simple in principle — call a different API. In practice, it means maintaining prompt compatibility across models with different context windows and system prompt behaviors, running evaluations so you know which model actually performs on your workloads, and building routing logic that accounts for latency, cost, and availability simultaneously. The complexity is real, which is why it needs to be solved at the platform level rather than reimplemented by every team that builds on AI.

A Decision You Are Making Now

The original argument for multi-provider AI was about quality: different models win at different tasks, and the winner keeps changing. Claude 3.5 Sonnet moved Dust's traffic from 20% Anthropic to 60% Anthropic within days of its release [10]. OpenAI's o3 solved reasoning problems that had stumped every model for months. Gemini Flash became the fastest low-latency option for summarization workloads. The leaders keep rotating.
The updated argument is more fundamental. Choosing a model is also choosing how exposed you are to one provider's infrastructure decisions, pricing strategy, and product roadmap. The model with the best benchmark score might be unavailable when you need it, repriced next quarter, or shut down because its compute budget got reallocated.
This is also a sovereignty question. When you put your organization's intelligence — its knowledge, its workflows, its decision-making — on a platform locked to a single model provider, you hand that provider control over your AI supply. They set the capacity, the price, and the priorities. In Europe, sovereignty also means being able to run on European models hosted in European infrastructure. In the US, it means not being captive to one provider's constraints. Either way, it means having a choice.
Organizations choosing AI platforms today are making infrastructure bets that take months to unwind. Porting an organization's workflows, prompts, and integrations off a single-model platform is a migration project, with all the cost and disruption that implies. The growth curves that created today's scarcity were on nobody's planning horizon eighteen months ago, and nobody can say with confidence when the next capacity crunch will hit or how long it will last.
Multi-provider AI is, at its core, insurance — against outages, against price hikes, against product cancellations, against a provider's roadmap diverging from yours. Like all insurance, it is cheapest to buy before you need it. The time to build in optionality is before the claim.
The founding bet at Dust — that building on any single provider was the wrong unit of analysis — was originally about intelligence. It is now also about supply.
Curious to see what multi-provider routing looks like for your team with Dust? Get in touch → 

​​[1] "Anthropic is paying SpaceX $15 billion per year," Axios, https://www.axios.com/2026/05/20/anthropic-spacex-compute, accessed June 2, 2026.
[2] "OpenAI shutters short-form video app Sora as company reels in costs," CNBC, https://www.cnbc.com/2026/03/24/openai-shutters-short-form-video-app-sora-as-company-reels-in-costs.html, accessed June 2, 2026.
[3] "NERC Issues Level 3 Alert, Reliability Guideline Focused on Large Load Challenges," NERC, https://www.nerc.com/newsroom/nerc-issues-level-3-alert-reliability-guideline-focused-on-large-load-challenges, accessed June 2, 2026.
[4] "Data to start your week: The AI capacity trap," Exponential View, https://www.exponentialview.co/p/data-to-start-your-week-the-ai-squeeze, accessed June 2, 2026.
[5] "Dario Amodei's 80x Growth Claim: What Anthropic's Q1 2026 Revenue Means," MindStudio, https://www.mindstudio.ai/blog/dario-amodei-80x-growth-anthropic-q1-2026-revenue/, accessed June 2, 2026.
[6] "Google reduces API rate limits for free tier," r/GeminiAI, Reddit, https://www.reddit.com/r/GeminiAI/comments/1pg4et5/google_reduces_api_rate_limits_for_free_tier/, accessed June 2, 2026.
[7] "Changes to GitHub Copilot Individual plans," GitHub Blog, https://github.blog/news-insights/company-news/changes-to-github-copilot-individual-plans/, accessed June 2, 2026.
[8] "Mistral AI CEO Mensch to French Lawmakers: Europe Has Two Years to Stop Losing the AI Race," French Tech Journal, https://www.frenchtechjournal.com/mistral-ai-ceo-mensch-to-french-lawmakers-europe-has-two-years-to-stop-losing-the-ai-race-before-the-race-is-over/, accessed June 2, 2026.
[9] "Claude loses its >99% uptime in Q1 2026," Hacker News, https://news.ycombinator.com/item?id=47543189, accessed June 2, 2026.
[10] "The Single-Model Trap: Why AI Platforms Need Multiple Providers," Dust Blog, https://dust.tt/blog/the-single-model-trap-why-ai-platforms-need-multiple-providers, accessed June 2, 2026.