← All posts / Industry

Sixty-Five Million a Year to Serve Free Models: Inside the Crusoe–Thinking Machines Inference Deal

Crusoe will serve Thinking Machines Lab's open models — Inkling, GLM 5.2 and 5.3 — on a dedicated HGX B200 cluster in a $65M/year deal that pushes Crusoe Managed Inference past $100M in contracted ARR less than a year after launch.

Sixty-Five Million a Year to Serve Free Models: Inside the Crusoe–Thinking Machines Inference Deal

There is a quiet irony at the heart of the AI infrastructure market: the models generating the most inference traffic are increasingly the ones nobody pays a license fee for. On September 23, 2026, that irony got a price tag. Crusoe, the vertically integrated “AI factory” company, and Thinking Machines Lab — the research lab founded by former OpenAI Chief Technology Officer Mira Murati — announced a $65 million annual agreement to run production inference for the lab’s open models on Crusoe Cloud.

The deal is a snapshot of where the AI stack’s money is actually flowing in late 2026: not into model licenses, but into the electricity, silicon, and serving infrastructure behind models that anyone can download for free.

What the deal actually covers

According to the announcement, Thinking Machines Lab will serve a range of production workloads through Crusoe Managed Inference — including its Inkling model family, the GLM 5.2 and 5.3 models, and the lab’s own fine-tuned variants. The hardware is a dedicated Tailored Deployment of NVIDIA HGX B200 systems, connected with NVIDIA Quantum-2 InfiniBand networking and “engineered for high-throughput, cost-efficient serving at scale.”

The Tailored Deployment framing matters. This is not a self-serve API or a shared serverless pool. Crusoe will run and support the cluster directly, providing what the company describes as a “dedicated, benchmarked, SLA-backed endpoint.” In other words, Thinking Machines gets the economics and headroom of owning an inference cluster without having to build, tune, and staff the serving stack itself.

The workload list is also telling. Inkling, released in July 2026, is a 975-billion-parameter mixture-of-experts model with roughly 41 billion parameters active per token — the strongest open-weight model to come out of a U.S. lab, and the flagship proof that Thinking Machines can ship frontier-class weights under a permissive license. GLM 5.2 and 5.3, meanwhile, are models the lab has adapted and served for its own products. Fine-tuned variants round out the picture: this is production traffic, not a benchmark run.

The economics of serving open weights

For Thinking Machines, the calculus is straightforward. Open-weight models do not generate license revenue — they generate usage. Every developer, agent framework, and enterprise that pulls the weights or hits the lab’s endpoints costs money to serve, and the cost scales with adoption, not with sales. A lab that wants to keep its models free, fast, and reliable at the frontier has to either become an infrastructure company or rent one.

“Crusoe Managed Inference took us from evaluation to production quickly and met the standards we hold our own systems to, giving us the scale and economics to reinvest in the research at the core of what we do,” said Myle Ott, ML Infra Lead at Thinking Machines Lab, in the announcement. The last clause is the strategic one: outsourcing the serving layer is what frees capital and engineering attention for research — the thing a lab led by Murati, and staffed heavily by ex-OpenAI alumni, is actually built to do.

There is also a second phase on the roadmap. The two companies say they are “looking to expand into batch inference for large-scale synthetic data generation, turning inference into a flywheel for research.” That is worth watching: batch synthetic-data generation is the hidden compute sink of modern post-training, and if Thinking Machines starts generating training data on Crusoe GPUs, a $65M inference contract could quietly become a much larger compute relationship.

A milestone for Crusoe’s inference business

For Crusoe, the deal is a proof point in a deliberate pivot. The company built its name on energy-first AI data centers — pairing gigawatt-scale power procurement with GPU fleets — and raised a $3.9 billion Series F on September 17, 2026 to double down on that vertically integrated platform. Managed Inference, which only reached general availability in November 2025, is the layer that turns owned data centers into a product a model lab can consume without thinking about hardware.

Thinking Machines Lab “joins a roster that has taken Crusoe Managed Inference past $100 million in contracted ARR less than a year from launch,” the company said. A single $65M contract does not make a business — and the ARR figure counts multi-year contracted value across the roster — but it does signal that managed inference has graduated from experiment to a real product line for the AI cloud providers. Rivals like Together AI, Fireworks, and the hyperscalers are chasing the same workloads, and open-model serving is becoming the segment where they compete most directly.

“Thinking Machines Lab has a very sophisticated engineering team that expects partners to meet their high bar as they grow,” said Erwan Menard, SVP of Product Management for Crusoe Cloud. “Our team of leading AI engineers built Crusoe Managed Inference to deliver cost-efficient, reliable infrastructure, inference, and model lifecycle tooling as-a-service.”

Why this matters beyond one contract

Stack the context up and the deal reads as a template for the open-weights economy. Thinking Machines is reportedly in talks to raise at around a $40 billion valuation while its models circulate freely; Crusoe is converting a $3.9 billion infrastructure war chest into recurring inference revenue. The money in open-weight AI does not vanish — it migrates down the stack, from licenses to tokens to electrons.

Three things follow. First, inference economics will increasingly decide which open models stay relevant: a model nobody can serve cheaply at scale is a research artifact, not a product. Second, the bonds between labs and infrastructure providers will keep tightening, because a dedicated SLA-backed cluster is a genuine moat around model quality of service. Third, watch the synthetic-data flywheel — if batch generation moves onto this cluster, the distinction between “inference deal” and “training deal” starts to blur, and Crusoe’s full-lifecycle pitch (it signed Perplexity to exactly such a train-and-serve agreement on September 15) becomes the standard shape of AI cloud contracts.

For now, the scoreboard is simple: $65 million a year, HGX B200s on InfiniBand, and the strongest open-weight model from a U.S. lab being served by a company whose thesis is that intelligence, like electricity, should be manufactured at industrial scale.