First Electron to Last Token: Crusoe Signs Perplexity to Its Full Model Lifecycle
Crusoe will train Perplexity's frontier models on dedicated GB300 NVL72 clusters and serve them through Managed Inference — while Crusoe's 1,800 employees standardize on Perplexity Enterprise Pro and Max.
The AI cloud market has a new reference deal for what “full lifecycle” actually means. On September 15, 2026, Crusoe — the self-styled “AI factory” company and one of the loudest advocates of vertically integrated, energy-first infrastructure — announced a multi-year partnership with Perplexity that covers both ends of the model pipeline: Perplexity will train its frontier models on Crusoe’s dedicated NVIDIA GB300 NVL72-powered clusters, and then serve those same models in production through Crusoe’s Managed Inference service.
What the deal actually covers
According to the announcement, datelined San Francisco, the agreement has two prongs. The first is compute: Perplexity’s full model lifecycle now runs on Crusoe Cloud, with training on dedicated GB300 NVL72 clusters linked by NVIDIA InfiniBand, and production serving handled by Crusoe Managed Inference — a service that only reached general availability in November 2025.
The second prong is the more unusual one. Crusoe will adopt Perplexity Enterprise Pro and Max for its roughly 1,800 employees, giving the infrastructure company’s staff access to Perplexity’s web and internal knowledge search, multi-step research, data analysis, and frontier model access. In other words, the compute vendor becomes the customer — a symmetry that both companies were happy to highlight.
“The fastest-moving AI companies need infrastructure that keeps pace across the entire model lifecycle and scales with them as they grow,” said Chase Lochmiller, co-founder and CEO of Crusoe. “Perplexity is building the future of how people find answers, and Crusoe is a key partner in making that possible, from the first electron to the last token.”
That closing phrase — “from the first electron to the last token” — is doing a lot of work in the press release. It compresses Crusoe’s core pitch: that owning the energy layer (power procurement and grid strategy), the data center layer, and the cloud software layer lets it optimize a workload end to end in a way that renting slices of someone else’s cloud does not.
The hardware underneath
The training side of the agreement runs on NVIDIA’s GB300 NVL72, the rack-scale Blackwell Ultra platform that has become the default building block for serious frontier training clusters in 2026. The specs explain why: each fully liquid-cooled rack integrates 72 Blackwell Ultra GPUs and 36 Arm-based Grace CPUs, with 130 TB/s of NVLink bandwidth, 20 TB of aggregate GPU memory offering up to 576 TB/s of bandwidth, and 2,592 Arm Neoverse V2 CPU cores per system. Every GPU gets 800 Gb/s of network connectivity through a ConnectX-8 SuperNIC, tying into Quantum-X800 InfiniBand or Spectrum-X Ethernet fabrics.
NVIDIA claims AI factories built on GB300 NVL72 deliver up to 50x the AI factory output of Hopper-based platforms — a figure built from a projected 10x improvement in user responsiveness (tokens per second per user) and 5x better throughput per megawatt, and one the company itself notes is subject to change. Even discounted, the platform is the current state of the art for the kind of dense, tightly-coupled training that frontier post-training runs demand.
Dion Harris, senior director of HPC and AI infrastructure solutions at NVIDIA, framed the deal as validation of that architecture: “Modern AI requires a unified computing architecture from research to deployment. Powered by NVIDIA GB300 NVL72 and NVIDIA InfiniBand, Crusoe Cloud enables Perplexity to train, fine-tune, and serve frontier models in production with the speed and efficiency agentic AI demands.”
Inference is the product
For Perplexity, the training half of the deal is table stakes; the inference half is existential. “Running AI at Perplexity’s scale means every millisecond of latency is felt by users,” said Aravind Srinivas, co-founder and CEO of Perplexity. “Crusoe is built around exactly that constraint. It’s high-throughput, low-latency inference backed by next-generation hardware.”
That is not marketing filler. A search-and-answer product serving millions of users in real time monetizes responsiveness directly — latency and throughput are not theoretical benchmarks for Perplexity, they are the product. Sourcing inference from a provider whose engine is explicitly optimized around those constraints, on the same platform where the models were trained, removes a whole class of handoff problems.
Crusoe’s Managed Inference service is powered by its proprietary inference engine built around MemoryAlloy, a cluster-wide key-value cache that eliminates duplicate prefills by letting GPUs fetch prefix caches from local and remote nodes. Crusoe has published benchmarks claiming up to 9.9x faster time-to-first-token and 5x higher throughput against vLLM on Llama-3.3-70B in a four-node deployment. The service comes in three flavors: serverless inference for spiky or experimental workloads, self-serve deployments now generally available, and tailored dedicated endpoints for proprietary models — and it is presumably the last of these that Perplexity’s production traffic will ride on.
Why this deal matters
Three larger storylines converge here.
The neocloud land grab moves up-stack. Crusoe’s rise has been one of the fastest in the infrastructure tier: a $600M Series D at a $2.8B valuation in December 2024, a $1.375B Series E above $10B in October 2025, and reportedly a $3B Series F at a $30B valuation in mid-2026 — with revenue projected to have grown from $276M in 2024 toward the $1B mark on the back of builds for customers like OpenAI’s Stargate program. Deals like this one, which anchor a named consumer-scale AI tenant across training and serving, are how a neocloud converts raw capacity into durable, multi-year revenue.
The training-to-serving gap is closing. For most of the AI boom, frontier labs trained on one platform and served on another — often a hyperscaler — accepting the operational seam in between. Perplexity collapsing its lifecycle onto a single platform is a bet that unified optimization (shared tooling, consistent networking, a caching layer that understands your models’ prefix patterns) beats best-of-breed assembly. If that bet is right, expect more labs to follow, and expect the hyperscalers’ bundling advantage to erode at the margin.
Compute-for-product swaps are becoming a thing. Crusoe standardizing its own workforce on Perplexity Enterprise Pro and Max mirrors a pattern of infrastructure deals increasingly packaged with reciprocal product adoption. It is a small gesture at 1,800 seats, but it signals how AI vendors now measure partnerships — not just in dollars but in becoming part of each other’s internal toolchain.
The deal’s financial terms were not disclosed, and multi-year infrastructure agreements in this market routinely carry committed-spend floors well into nine figures. What is disclosed is the shape: one platform, two directions of value flow, and a very explicit claim that the future of AI infrastructure competition is lifecycle coverage rather than raw GPU inventory.
For an industry still digesting what it means that inference — not training — is now the dominant cost center of deployed AI, Crusoe and Perplexity have offered one concrete answer: own the loop from the first electron to the last token.
Sources
- [1] https://www.crusoe.ai/resources/newsroom/crusoe-perplexity-partnership
- [2] https://www.unite.ai/crusoe-signs-multi-year-deal-to-power-perplexity-training-and-inference/
- [3] https://www.hpcwire.com/aiwire/2026/09/16/crusoe-and-perplexity-announce-multi-year-partnership-across-full-model-lifecycle/
- [4] https://www.lightreading.com/ai-machine-learning/crusoe-and-perplexity-announce-multi-year-partnership