IBM and Together AI Bet $240 Million on NVIDIA B300: A New Front in the Open-Source AI Infrastructure War
IBM signed a $240M deal with Together AI to deploy a large NVIDIA HGX B300 inference cluster on IBM Cloud — positioning Big Blue as the enterprise host for the open-source AI revolution.
On August 11, 2026, IBM announced a multi-year, $240 million agreement with Together AI — a deal that will see a large cluster of NVIDIA HGX B300 systems deployed on IBM Cloud, dedicated to open-source AI model inference. The cluster is expected to come online in the first quarter of 2027, and it represents a significant strategic pivot for both companies: IBM is staking its claim as the enterprise-grade host for the surging open-source AI ecosystem, while Together AI gains the infrastructure scale and enterprise credibility it needs to compete with hyperscaler-backed platforms.
The Deal in Detail
Under the agreement, IBM will provision a dedicated fleet of NVIDIA HGX B300 systems — the Blackwell Ultra generation — connected via NVIDIA Spectrum-X Ethernet networking. Together AI will use this capacity to serve inference workloads for hundreds of open-source models to its enterprise customers, who increasingly rely on open weights like Meta’s Muse Glimmer, Alibaba’s Qwen family, and DeepSeek’s V4-Flash to power production applications without the per-token costs and vendor lock-in associated with proprietary APIs.
The financial structure is straightforward: Together AI commits to $240 million over a multi-year period, and IBM provides the physical infrastructure, cloud orchestration layer, and enterprise compliance stack. Together AI retains responsibility for the model-serving software, developer experience, and customer relationships. For IBM, this is essentially a high-value anchor tenant for its GPU cloud business — a signal that Big Blue is serious about competing for AI compute workloads alongside AWS, Google Cloud, and Microsoft Azure.
Why the NVIDIA HGX B300 Matters
The choice of NVIDIA’s HGX B300 platform is central to understanding why this deal is more than a routine cloud procurement. The B300 — powered by the Blackwell Ultra GPU — represents NVIDIA’s most advanced data-center silicon as of mid-2026. Each HGX B300 system packs eight Blackwell Ultra SXM GPUs with a combined 2.3 TB of HBM3e memory and up to 144 petaFLOPS of FP4 tensor compute. That is roughly 11 times faster inference and 4 times faster training compared to the previous Hopper generation, according to NVIDIA’s own benchmarks.
For Together AI’s use case — high-throughput inference of large language models — the B300’s massive memory pool and bandwidth are the critical advantages. Models in the 30B–70B parameter range, which once required multi-GPU sharding with significant latency penalties, can now run on a single 8-GPU node with headroom to spare. This translates directly into lower cost per token, faster response times, and the ability to serve more concurrent users per physical server. The Spectrum-X Ethernet fabric, which provides 400 Gb/s per GPU with adaptive routing, ensures that multi-node deployments scale without the bottleneck that plagued earlier InfiniBand-dominated architectures.
In practical terms, the B300 cluster will allow Together AI to offer inference for frontier-scale open models — potentially including 100B+ parameter architectures — at price points competitive with, or cheaper than, proprietary API providers. That is the thesis driving the entire deal: open-source AI inference is becoming a viable enterprise alternative, and the infrastructure to support it at scale is now being built.
The Open-Source AI Inference Market
To understand why IBM and Together AI are making this bet, it is essential to look at the explosive growth of the open-source AI model ecosystem. According to data cited by Reuters and CNBC, Chinese open-weight models alone now account for 41% of downloads on Hugging Face and 61% of tokens processed on OpenRouter — the leading AI model gateway. DeepSeek’s V4-Flash model processed 7.22 trillion tokens in a single week, claiming the number-one global spot on OpenRouter. Meta’s open-weight strategy, championed by CEO Mark Zuckerberg in a sweeping 6,500-word manifesto, has further accelerated adoption.
This shift creates enormous demand for inference infrastructure that is not tied to a single proprietary model vendor. Enterprises want choice: the ability to switch between DeepSeek, Qwen, Llama, or specialized models without re-architecting their applications. Together AI has positioned itself as the neutral inference layer — offering 200+ open-source models through a single API — and the IBM deal gives it the compute capacity to back that promise at enterprise scale.
Together AI itself has been on a remarkable trajectory. Founded in 2022 by serial entrepreneur Vipul Ved Prakash and a team of academic researchers, the company raised $1 billion at a $7.5 billion valuation in early 2026, driven by the conviction that open-source models would eventually dominate AI usage. The NYT reported in July 2026 that Together AI’s revenue was surging as companies sought cheaper alternatives to OpenAI and Anthropic. The IBM partnership is the clearest sign yet that this conviction is being validated by the market.
IBM’s Strategic Calculus
For IBM, the deal is about reclaiming relevance in the AI infrastructure race. Despite its decades of enterprise IT dominance, IBM Cloud has been a marginal player in the AI compute market, dwarfed by AWS (which controls a significant share of the neocloud GPUaaS market), Microsoft Azure (which has bet heavily on OpenAI integration), and Google Cloud (which leverages its TPU advantage). By partnering with Together AI, IBM gains several things at once.
First, it gets a high-profile customer that validates IBM Cloud’s GPU capabilities for demanding AI workloads. Second, it positions IBM as the enterprise-compliant alternative to smaller neoclouds like CoreWeave, Lambda Labs, and Voltage Park — companies that have the GPUs but lack IBM’s security certifications, hybrid-cloud integration, and global enterprise sales force. Third, it deepens IBM’s relationship with NVIDIA at a time when GPU supply remains the binding constraint in the AI industry. IBM is effectively leveraging its enterprise “street cred” — as one industry analyst put it — to win infrastructure deals that pure-play neoclouds cannot easily close.
The deal also fits into IBM’s broader AI strategy, which centers on its watsonx platform, Red Hat OpenShift AI, and consulting services. By hosting Together AI’s inference cluster, IBM creates a natural funnel: enterprises that come for open-source model inference may stay for IBM’s full-stack AI governance, compliance, and integration offerings. It is a classic IBM land-and-expand play, updated for the generative AI era.
Competitive Implications
The IBM–Together AI partnership sends ripples across the AI infrastructure landscape. For Microsoft, it highlights a growing tension: while Azure has benefited enormously from its exclusive OpenAI partnership, the open-source model ecosystem represents a divergent and potentially larger market that Microsoft’s own infrastructure strategy has under-served. Microsoft’s recent moves to build in-house AI models (MAI series) and its own Maia silicon suggest an awareness of this gap, but the company has not yet matched the open-model-first positioning of Together AI.
For AWS, the deal represents competition in its core AI inference business. AWS has invested heavily in its own Inferentia and Trainium chips and has positioned Bedrock as a multi-model platform, but Together AI’s pure-play focus on open weights — combined with IBM’s enterprise reach — creates a credible alternative channel for enterprise inference workloads.
For the neoclouds — CoreWeave, Lambda, Voltage Park, and others — the message is that the bar for enterprise credibility is rising. Raw GPU capacity is becoming commoditized; the differentiators are shifting toward compliance, security, networking quality, and the ability to integrate with enterprise IT environments. IBM’s entry into this space, via the Together AI partnership, raises the competitive stakes for everyone.
Looking Ahead
The cluster is not expected to be available until Q1 2027, which means the real test of this partnership lies ahead. Key questions remain: Will enterprises actually migrate inference workloads from proprietary APIs to open-source alternatives hosted on IBM Cloud? Will Together AI’s model-serving stack achieve the latency and reliability that production applications demand? And will NVIDIA’s B300 supply constraints — which have plagued the industry throughout 2026 — allow IBM to actually deliver the promised capacity on schedule?
What is clear is that the open-source AI inference market has crossed an inflection point. When a company of IBM’s stature commits nearly a quarter-billion dollars to host open models on the most advanced GPU hardware available, it signals that open-source AI is no longer a developer novelty — it is becoming the enterprise default. The next eighteen months will reveal whether IBM and Together AI have timed this bet correctly.
Sources
- [1] https://newsroom.ibm.com/2026-08-11-IBM-and-Together-AI-Sign-Multi-Year-Agreement-to-Scale-Open-Source-AI-Inference-with-NVIDIA-AI-Infrastructure-on-IBM-Cloud
- [2] https://www.reuters.com/business/ibm-together-ai-ink-240-million-deal-nvidia-powered-ai-inference-cluster-2026-08-11/
- [3] https://www.techzine.eu/news/infrastructure/143545/ibm-builds-240-million-inference-cluster/
- [4] https://finance.yahoo.com/technology/ai/articles/ibm-together-ai-sign-240-134324558.html
- [5] https://www.storagereview.com/news/ibm-and-together-ai-put-240m-into-a-dedicated-hgx-b300-inference-cluster-on-ibm-cloud
- [6] https://www.morningstar.com/news/dow-jones/202608115403/ibm-together-ai-ink-240-million-agreement-to-scale-ai-deployments
- [7] https://www.fiercewireless.com/cloud/ibm-lends-together-ai-its-enterprise-street-cred-240m-compute-deal
- [8] https://www.tradingview.com/news/stocktwits:086d07fea094b:0-ibm-s-240m-together-ai-deal-for-nvidia-systems-puts-it-in-the-neocloud-business/