From 50% to Minus 50%: The Gross-Margin Collapse Pushing Harvey, Abridge and Ramp Into Open Weights
A Bloomberg investigation details how frontier-model API costs crushed the unit economics of top AI application startups — Harvey's gross margin fell from 50% to minus 50% in six months — and why open weights and in-house models are now the survival plan.
When Your Best Customers Become Your Biggest Cost
The most important number in the AI application business this month is not a benchmark score. It is a gross margin: minus 50 percent.
According to a Bloomberg report published September 21, 2026, Harvey — the $15.6 billion legal-AI startup that built its business on top of OpenAI’s and Anthropic’s models — watched its gross margin collapse from roughly 50% at the start of 2026 to negative 50% by June, as customer usage of its AI agents spiked. Every dollar of revenue was accompanied by more than a dollar of inference cost. The company was, in effect, paying its users to grow.
Harvey is not alone. Bloomberg identifies a broader pattern across the application layer: Abridge, the ambient clinical-documentation company whose AI now sits in on doctor visits across major US health systems; Ramp, the spend-management platform that raised $750 million in June; and Rogo, the AI analyst used on Wall Street — all are embracing open-weight models or training their own to reduce what the frontier labs charge them every month. The report lands amid what Techmeme tracked as one of the biggest stories of the week: keeping up with the AI frontier now costs more than most companies will ever raise.
The Math That Broke the Application Layer
To understand why a 100-point margin swing is possible in six months, look at how the economics of agentic AI differ from the SaaS playbooks these companies were valued on.
A traditional software company serves an incremental customer for nearly nothing; gross margins of 70–80% are the norm. An AI application company instead resells compute. Every contract its lawyer-AI reads, every doctor visit its scribe transcribes, every pitch deck its banking agent drafts triggers tokens billed at frontier-lab prices. When usage is low, the wrapper looks like great software. When usage spikes — exactly what every one of these companies has been celebrating in their funding announcements — the cost curve goes vertical.
Harvey’s case is the cleanest illustration. The company crossed $350 million in annualized revenue in 2026 on the strength of agents that run multi-hour legal workflows. But long-horizon agent tasks consume orders of magnitude more tokens than a chat query, and the more successful the product became, the deeper the losses on each account. According to a person familiar with the matter cited by Bloomberg, gross margins fell from about 50% at the start of the year to minus 50% by June — a swing that would be catastrophic for a public company and is merely existential for a private one.
There is a second, quieter factor: pricing power. Frontier-model prices have fallen, but the frontier keeps moving. The Bloomberg report notes that Anthropic recently signed a $13.7 billion compute deal to keep pace — a scale of spending that gets recapitulated in the API prices everyone downstream pays. Application startups funding frontier R&D through their inference bills is a business model with no floor.
Open Weights as the Escape Hatch
The response, across the board, is to stop renting intelligence.
Harvey made the most aggressive move. On August 20 it released Harvey Tenet, its first in-house model — not trained from scratch, but post-trained on Kimi K3, the open-weight model from China’s Moonshot AI, working with Fireworks AI on a corpus of legal work product. On Harvey’s own evaluations, Tenet matches or beats the frontier models it replaces on long-horizon legal tasks, at a fraction of the per-token cost. A month later, the company closed a $550 million round at a $15.6 billion valuation explicitly to fund this model-building agenda — raising money from investors, including OpenAI itself, to build an alternative to OpenAI’s economics.
Abridge took the partnership route. At its June platform keynote, the company announced it will train and fine-tune on top of Nvidia’s open-weights Nemotron family, using its own de-identified clinical data. For a company whose product is literally listening to millions of doctor visits, proprietary clinical data plus an open base model is a defensible stack that no API price hike can repossess.
Ramp and Rogo round out the list. Ramp’s June raise — $750 million — was framed by observers as partly a war chest for infrastructure independence. Rogo, the $2 billion AI-analyst startup now backed strategically by Barclays, BNP Paribas and other banks it serves, has been steadily moving workflow onto its own post-trained models. The Bloomberg framing is blunt: these companies are embracing open weights or training their own models to reduce expensive reliance on frontier labs.
Why This Reshapes the Whole Stack
Three implications are worth taking seriously.
First, the wrapper backlash is becoming an exodus. For two years the critique of application-layer AI companies was that they were “thin wrappers” with no moat. The 2026 answer — post-train an open-weight base on your proprietary workflow data — converts that critique into an architecture. The moat is no longer the model; it is the data and the post-training pipeline. Companies like Harvey and Abridge are betting their valuations on it.
Second, the frontier labs are losing their best customers. Harvey, Abridge, Ramp and Rogo are precisely the high-volume, high-profile accounts that funded the frontier labs’ own growth. OpenAI invested in Harvey; now Harvey trains on Moonshot’s weights. Every flagship customer that internalizes its inference is a structural headwind to the labs’ revenue — one of the reasons OpenAI’s own IPO timing (now ruled out for 2026) and Anthropic’s margin discipline are under such intense scrutiny.
Third, open weights are now a boardroom cost decision, not an ideology. The open-versus-closed debate used to be about research freedom and safety. In 2026 the deciding argument is a gross-margin line item. Notably, several of the open-weight bases being adopted — Kimi K3, Nemotron — are either Chinese or Nvidia-trained, which folds this economics story into the geopolitical one: US application companies are quietly building on Chinese open weights because the arithmetic demands it.
The Counterargument
The frontier labs’ case is that open-weight models trail the frontier on the hardest tasks, and that vertical post-training cannot close gaps in reasoning depth. Mozilla’s September report estimated the gap at roughly four months — a head start, not a moat. If the frontier labs are right that capability gaps compound, today’s minus-50% margins could look like a cheap tuition. But four months is also roughly the cadence at which Harvey now ships post-trained models, and vertical data does not trail anything.
What to Watch
- Whether Harvey’s Tenet graduates from research preview to default engine for production legal work
- Gross-margin disclosures in the next funding rounds of Abridge, Ramp and Rogo — the market will demand them
- Frontier-lab API price cuts timed to retain flagship accounts, and whether they arrive before or after more exits
- The next $10B+ compute deal and how much of it gets passed downstream
The application layer has stopped asking whether open weights are good enough. It has started asking whether anything else is affordable.
Sources are listed in the article metadata.
Sources
- [1] https://www.bloomberg.com/news/articles/2026-09-09/legal-ai-startup-harvey-hits-15-6-billion-value-with-550-million-round
- [2] https://thenextweb.com/news/model-costs-startups-open-weights
- [3] https://www.harvey.ai/blog/post-training-update-harvey-tenet
- [4] https://www.healthcare-brew.com/stories/abridge-nvidia-eli-lilly-ai-clinical-platform
- [5] https://www.fiercehealthcare.com/ai-and-machine-learning/nvidia-abridge-collaborate-develop-healthcare-specific-ai-model
- [6] https://www.wsj.com/finance/banking/wall-streets-favorite-ai-startup-sets-its-sights-on-wealth-management-61b8b5d1