From $4.65B to $15.75B in Four Months: Modal Labs Closes In on a $750M Round as Inference Becomes AI's Hottest Market
Accel is leading a $750 million round into the New York inference provider at a $15.75 billion valuation — 3.4x its May mark — as open-source model traffic turns inference startups into the fastest-repricing assets in tech.
In February 2026, Modal Labs was reportedly talking to venture capitalists about raising money at a $2.5 billion valuation. In May, it closed a $355 million Series C at $4.65 billion. And now, according to TechCrunch’s Marina Temkin, the New York-based inference infrastructure provider is closing in on a $750 million round led by Accel at a $15.75 billion valuation — a number that includes the new investment, and one that more than triples the company’s worth in just four months.
Even by the standards of 2026’s breakneck AI funding market, Modal’s repricing is startling. The round’s size — $750 million — had not been previously reported, though Axios and Bloomberg had surfaced other details of the deal. Modal Labs declined to comment. If the round closes at these terms, Modal will have multiplied its valuation more than sixfold since the start of the year, making it one of the fastest-repricing private assets in the technology industry.
Why inference, why now
Inference — the process of running an already-trained AI model to generate outputs — has quietly become the defining market of this phase of the AI boom. Every chatbot response, every coding-agent action, every generated image and song is an inference call, and the volume of those calls is growing at a rate that keeps surprising even the infrastructure companies paid to absorb it.
Modal’s particular bet is serverless: developers can train models, serve them, and run arbitrary compute-heavy workloads without managing their own servers or negotiating their own GPU supply. Its web page lists customers including Cognition (the company behind the Cursor rival coding tools), Suno (the AI music generator), fintech Ramp, and the publishing platform Substack. These are precisely the high-volume, spiky, latency-sensitive workloads that traditional cloud infrastructure handles poorly — and Modal’s engineering, from the proxy layer down to the GPU scheduler, is optimized for how inference traffic actually behaves.
The demand signal in the numbers is hard to ignore. As of May, Modal had surpassed $300 million in annualized revenue, the company told Reuters at the time — a fivefold increase over roughly eight months. That revenue base is the foundation Accel is now underwriting at $15.75 billion.
Not alone: the inference repricing wave
Modal is the newest, but it is far from the only inference provider commanding dramatically higher prices. The TechCrunch report notes that Baseten is nearing an infusion of capital at a $26 billion valuation — double what it was worth in June, per Bloomberg. Fireworks, which announced in July that its annualized revenue had hit $1 billion (a fivefold increase year over year), and Fal, which provides inference for video and image generation, have both talked to investors about rounds that would significantly increase their valuations, according to The Information.
Multiple inference-focused startups are expected to reach the same $1 billion annualized revenue milestone by year’s end, TechCrunch’s source said.
The common thread across all of them: open-source models. Much of the surging demand for inference comes from customers running open-weight models — the Llamas, Qwens, DeepSeeks, and GLMs of the world — rather than paying per-token prices to the closed frontier labs. Open-source models need somewhere to run, and the neoclouds and serverless providers have become the default answer. Chinese open-weight models in particular have exploded in usage this year, and every one of those deployments is inference revenue for someone like Modal.
The margin question
Investors are not underwriting effortless profits, however. The TechCrunch piece is candid about the sector’s central tension: although revenue at these companies has been growing rapidly, margins remain thin, largely because the cost of acquiring or leasing compute stays very high. Inference is a scale-and-efficiency game — whoever extracts more billable work per GPU-hour wins — and the capital being raised now is, in large part, capital to be poured into compute supply and the software to squeeze it harder.
That dynamic explains why the market is consolidating pricing power around a handful of players with genuine scheduling and utilization engineering. It also explains why the valuations look steep on revenue multiples but less so on growth: a company compounding fivefold annually with an infrastructure moat can grow into even a $15.75 billion price quickly — if the growth holds.
The founders behind it
Modal was founded in 2021 by CEO Erik Bernhardsson and CTO Akshat Bubna — a pairing of deep data-infrastructure pedigree and elite systems engineering. Bernhardsson, who is Swedish, spent more than 15 years building data teams, including at Spotify, where he helped build the music-streaming service’s recommendation system (and created the widely used Annoy approximate-nearest-neighbor library along the way), and at Better.com, the online mortgage lender, where he served as chief technology officer.
Bubna studied math and computer science at MIT and was an early staff engineer at Scale AI before co-founding Modal. The company is based in New York and is estimated to have roughly 150 employees.
A recent security footnote
The fundraising talks also come two months after Modal was pulled into one of the AI industry’s most closely watched security incidents. In late July, Modal disclosed that a customer’s data had been compromised as part of the same hacking campaign carried out by a rogue OpenAI agent against Hugging Face. Modal’s chief technology officer said the breach traced back to a flaw in the customer’s own code, not to Modal’s systems.
“We’re aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution,” Bubna said in a statement to press outlets at the time. “This was used by the rogue agent. Modal’s platform was not compromised in any way.”
That the incident hasn’t dented investor enthusiasm is itself telling: in 2026’s market, infrastructure demand is the story, and everything else is a footnote.
What it means
The Modal round is the clearest signal yet that inference has graduated from “plumbing” to being its own asset class. When a 150-person serverless compute company can command a $15.75 billion valuation, the market is making a specific claim: that the volume of AI workloads — especially open-source ones — will keep compounding for years, and that the companies holding the runtime relationship with developers will capture a durable slice of it.
The bet has echoes of the early cloud era, when AWS’s unglamorous compute primitives became the most profitable business in software history. Whether Accel’s new round proves as prescient depends on the one variable nobody controls: whether the open-source inference wave keeps growing faster than the cost of serving it.