One Workload, Six Kinds of Silicon: Gimlet Labs Raises $300M at a $3B Valuation to Orchestrate AI's Heterogeneous Future
Six months after an $80M Series A, Gimlet Labs has raised $300M led by a16z at a $3B valuation, with Arm and Microsoft's M12 joining, to run AI workloads across NVIDIA, AMD, Intel, ARM, Cerebras, and d-Matrix chips.
The strangest thing about the AI boom, according to Andreessen Horowitz general partner Raghu Raghuram, is that it has been built largely on a single technology stack anchored by one company’s chips. “The history of computing is never that way,” he told Bloomberg this week — “especially a market that’s going to be trillions and trillions of dollars and so many different use cases.” That conviction is now worth $3 billion on paper.
Gimlet Labs, a San Francisco startup that helps customers divide AI tasks between multiple types of chips, has raised $300 million in a round led by Andreessen Horowitz, vaulting its valuation to $3 billion just six months after an $80 million Series A. The round, confirmed by CEO Zain Asgar in an interview, includes new investments from Arm Holdings and M12, Microsoft’s venture fund — two backers with an obvious strategic interest in a world where AI workloads no longer default to a single vendor.
What Gimlet actually does
At its core, Gimlet builds what it calls the industry’s first “multi-silicon inference cloud” — a software layer that takes an AI workload, breaks it into pieces, and maps each piece to whichever chip can run it best. The platform orchestrates across hardware from NVIDIA, AMD, Intel, Arm, Cerebras, and d-Matrix, spanning conventional GPUs alongside SRAM-centric and other exotic silicon.
The founding insight is that modern AI applications — especially agentic ones — are not monolithic. An agent might chain multiple models and modalities together: a small fast model for planning, a large frontier model for reasoning, specialized accelerators for retrieval or decoding. Running that entire pipeline on one homogeneous fleet of GPUs is convenient but wasteful. Different chips have wildly different economics for different operations, and the deltas are large enough that orchestration itself becomes a product.
Gimlet started by focusing on one very specific task — running AI models superfast — because that is where combining different chip types is yielding the most compelling results, Asgar said. The company emerged from stealth in late 2025 already claiming eight-figure revenues, an unusually strong position for a startup that had raised only a $12 million seed backed by venture firm Factory and angels including Intel CEO Lip-Bu Tan, Figma head Dylan Field, and Raghuram himself.
Peeling the onion into the data centre
But the more interesting part of the story is how the company’s scope has expanded. As Gimlet worked to deploy its software, it hit a wall that had nothing to do with algorithms: customers did not know how to arrange data centre hardware for complicated multi-chip setups.
“We figured out all this really cool tech around how to distribute workloads to different types of chips,” Asgar said. “But one of the challenges that we realized is that nobody builds data centres in a heterogeneous manner. So we kind of started peeling the onion one layer at a time.”
Different chip types often need different cooling systems or operating temperatures — a rack designed around GPU airflow is not automatically friendly to silicon with a different thermal profile. So Gimlet now helps customers set up their data centres, and is even working on data centres of its own. The startup that began as a scheduling layer has become an infrastructure company.
That evolution explains both the speed and the size of the round. Asgar said the fundraising came together rapidly because the company received several unsolicited term sheets from interested investors. When a startup with eight-figure revenue and a working product sits at the intersection of the two hottest themes in AI infrastructure — heterogeneous compute and inference economics — investors do not wait for a process.
Why heterogeneity is suddenly the consensus
For most of the past three years, “AI infrastructure” effectively meant a vertically integrated stack: one dominant chip vendor, its software ecosystem, and hyperscalers building around it. That monoculture is now breaking down for three reasons.
First, supply. Demand for top-tier accelerators still outstrips supply, and enterprises that cannot get enough of one chip are increasingly willing to mix. Second, cost. Inference — the phase where trained models actually serve users — is becoming the dominant share of AI compute spend, and it is far more price-sensitive than training. Novel silicon like Cerebras’s wafer-scale engines or d-Matrix’s in-memory compute can deliver dramatically better cost-per-token for specific operations, but only if workloads can actually reach them. Third, sovereignty and supply-chain resilience: governments and enterprises alike want multi-vendor optionality.
Gimlet is not alone in spotting this. Last month, a newer rival called Callosum raised $100 million in early financing from backers including the UK’s public AI fund, for software that routes AI work to whichever chip and model can run it most cheaply. Cambridge-based Callosum claims 4x speed and 70% lower compute costs. The venture market is clearly pricing in that the orchestration layer above the silicon becomes enormously valuable if the silicon itself fragments.
The strategic subtext of the cap table
The identity of the new investors matters as much as the money. Arm Holdings, whose instruction set architecture underpins the energy-efficient CPU and accelerator designs proliferating in AI data centres, is now both a backer and a collaborator — Gimlet is working with Arm to make its software compatible with various forms of Arm chip technology. For Arm, investing in the orchestration layer that routes workloads to the best chip is a way to ensure “best” is judged fairly across architectures.
M12’s participation extends the same logic to Microsoft, which has its own reasons to hedge compute strategy across vendors as it balances custom silicon, merchant GPUs, and a sprawling AI cloud. And a16z doubling down — Raghuram led the seed and is leading this round — signals conviction that “the world will be multi-silicon,” in his words, with “multiple different types of workloads with very, very different demands.”
What to watch
Gimlet’s customers include AI labs and financial services companies, though Asgar declined to name them. The open questions are the usual ones for an infrastructure startup riding a thesis: can the software layer stay neutral and trusted when its investors include chip companies with their own interests? Will chip vendors build or bundle comparable routing capability natively, commoditizing the layer? And can a company that must understand six silicon architectures deeply move fast enough against rivals optimizing for just one?
What is no longer debatable is the direction. The era of AI running on one kind of chip is ending — not because NVIDIA stumbled, but because a market measured in trillions of dollars was never going to stay homogeneous. Gimlet’s $3 billion valuation is an early, oversized bet that the operating system for that fragmented future is a company worth owning.
Sources
- [1] https://www.theedgemarkets.com/node/816853
- [2] https://www.techmeme.com/260903/p53
- [3] https://siliconangle.com/2026/03/23/multi-chip-inference-cloud-startup-gimlet-labs-receives-80m-solve-one-ais-biggest-bottlenecks/
- [4] https://techcrunch.com/2026/03/23/startup-gimlet-labs-is-solving-the-ai-inference-bottleneck-in-a-surprisingly-elegant-way/
- [5] https://gimletlabs.ai/blog/announcing-series-a