← All posts / Industry

No HBM, No Problem: Positron Closes $875M at $5B to Ship Its Memory-First Asimov Chip

The Reno AI-inference startup's Series C was co-led by NEA, Atreides and Valor; its Asimov accelerator swaps scarce HBM for cheap DDR5 and claims Nvidia-class inference at a fraction of the power.

No HBM, No Problem: Positron Closes $875M at $5B to Ship Its Memory-First Asimov Chip

Reno, Nevada is not where most people expect America’s next serious challenge to Nvidia’s data-center dominance to come from. But Positron AI, the inference-hardware startup founded in 2023 by Thomas Sohmers and Mitesh Agrawal, just closed a $875 million Series C at a $5 billion post-money valuation — a round big enough to fund the tape-out, production ramp, and market entry of its next-generation accelerator, Asimov.

The round, announced September 10, 2026, was co-led by NEA, Atreides Management, and Valor Equity Partners, according to the company’s press release. It caps a remarkable twelve months for the company: a $50M+ Series A in June 2025, a $230M Series B at a $1B+ valuation in February 2026 led by financial-trading firms including backing from Arm and Qatar’s investment fund, and now a Series C that values the business at five times its February mark. Total funding now exceeds $1.15 billion.

Why investors are betting against HBM

To understand why a 24-person-team-turned-chip-company keeps raising nine-figure rounds, you have to understand the bet: the economics of AI inference are broken, and HBM is the reason.

Modern AI accelerators from Nvidia, AMD, and others depend on high-bandwidth memory (HBM) stacked next to the compute die. HBM delivers enormous bandwidth — Nvidia’s newly announced Rubin GPUs pack 288GB of HBM4 good for 22 TB/s — but it is expensive, supply-constrained, and power-hungry. During inference (as opposed to training), large language models are largely memory-bound: the chip spends most of its time waiting for model weights to arrive from memory, not doing math. A compute-dense, HBM-laden GPU is a poor fit for that workload profile.

Positron’s answer is a “memory-first” architecture. Asimov is designed to support 2TB+ of memory per accelerator — 864GB to 2.3TB of realizable capacity — using conventional DDR5 DRAM, the same commodity memory found in laptops and servers, instead of HBM. As The Register reported when the Series B closed, Asimov trades peak bandwidth (topping out around a third of Rubin’s) for sheer capacity and cost efficiency. When your workload is bound by how much of the model fits in memory and how cheaply you can serve it, 2.3TB of DDR5 beats 288GB of HBM4 on dollars-per-token.

Four Asimov chips form a Titan system with 8TB+ of memory, with production targeted for early 2027 and tape-out in late 2026 — timelines this new round is explicitly designed to fund.

The Atlas track record

Asimov is not a paper architecture. Its predecessor, Atlas, has been shipping to customers since 2025 and provides the credibility behind the raise:

  • 3.5× better performance per dollar and up to 66% lower power usage than Nvidia’s H100, per company claims
  • Roughly 93% hardware utilization on transformer workloads — versus typical GPU utilization that often sits far lower on the same jobs
  • Runs models from the Hugging Face transformers library with zero code changes
  • Early traction among hyperscalers and financial-trading firms, whose latency-sensitive inference workloads reward the architecture’s strengths

Those numbers are company-provided and deserve scrutiny — but the customer mix matters. Trading firms were lead investors in the Series B precisely because they run inference at scale and could verify the unit economics with their own money on the line.

The context: an inference gold rush

The Series C lands amid a broad repricing of inference-focused silicon. OpenAI confirmed this week it is co-developing next-generation AI chips with Samsung. Analog Devices is buying edge-AI chipmaker Alif for $1.35B. Kepler Computing just exited stealth with $468M for ferroelectric-memory compute. Chinese AI chipmakers are raising prices 20–50% on HBM shortages. And Nvidia itself is lining up two gigawatts of AI factories in Australia.

Within that frenzy, Positron occupies a distinctive niche: it is not trying to out-engineer Nvidia at the frontier of process nodes and HBM stacks. It is attacking the cost structure of inference — betting that as AI moves from novelty to infrastructure, the winners will be those who can serve a token at the lowest energy and capital cost, using commodity components nobody else is fighting over.

That thesis has a second-order effect worth watching: if DDR5-based inference accelerators prove out at scale, they relieve pressure on the HBM supply chain that is currently throttling the entire industry — including for Chinese AI firms facing export controls and 20–50% price hikes on memory-bound designs.

What could go wrong

Skepticism is warranted on three fronts. First, Rubin-class GPUs will not stand still — Nvidia’s roadmap pushes HBM4 capacity and bandwidth higher each generation, and its CUDA software moat remains the industry default. Second, DDR5 bandwidth is a real constraint: Asimov’s ~3 TB/s ceiling versus Rubin’s 22 TB/s means workloads that are genuinely bandwidth-hungry (long-context, high-batch inference) may still favor GPUs, and Positron must win on cost-per-token in the segments it targets rather than raw speed. Third, 2027 production is a long time away in a market moving this fast — between tape-out and shipment, two more Nvidia generations will arrive.

But $875 million buys a lot of conviction. The round signals that top-tier investors believe the inference market is large enough — and Nvidia’s pricing power resented enough — to sustain a serious second source. Whether Asimov tape-outs on schedule and delivers the promised economics will be one of the more consequential hardware stories of 2027.

For now, a company that started in Reno with a contrarian thesis — that the future of AI inference looks more like cheap, capacious server memory than exotic stacked dies — has the balance sheet to test it at scale.