← All posts / Models

Not a Wrapper, a Foundation: Salesforce's Koa Bakes 27 Years of CRM Into Its Own Reasoning Model

Built by post-training NVIDIA's open-weight Nemotron 3 Super 120B with GRPO, Koa is Salesforce's first proprietary reasoning model — cutting CRM task errors threefold while keeping every weight inside its own trust boundary.

Not a Wrapper, a Foundation: Salesforce's Koa Bakes 27 Years of CRM Into Its Own Reasoning Model

At Dreamforce 2026 in San Francisco — which opened its doors today, September 15 — Salesforce and NVIDIA announced Koa, Salesforce’s first proprietary CRM reasoning model. It is the clearest signal yet that the “no wrappers” era of enterprise AI has arrived: instead of renting reasoning from a frontier lab, Salesforce post-trained NVIDIA’s open-weight Nemotron 3 Super 120B foundation model into a specialist that understands how businesses actually run — the structure of a deal, the lifecycle of a service case, the workflows that vary across industries.

Marc Benioff’s framing is blunt: “The most most valuable thing Salesforce has built isn’t our platform — it’s the accumulated knowledge of how enterprise business actually works. With Koa, that knowledge is put inside the model itself.” Jensen Huang counters from the silicon side: open Nemotron models give Salesforce “the foundation to turn decades of enterprise expertise into specialized AI.”

What Koa actually is

Koa is not a fine-tune bolted onto a chat model. The technical report, published simultaneously on arXiv (2609.15066) by a 25-author Salesforce AI Research team, describes an enterprise language model built by post-training Nemotron-3-Super-120B with reinforcement learning using Group Relative Policy Optimization (GRPO) — the same family of RL algorithms that reshaped reasoning-model training after DeepSeek-R1.

Two details stand out from the paper:

  • No customer data was used. The training corpus was built entirely from synthetic scenarios simulating enterprise workflows across more than 14 industries — manufacturing, financial services, healthcare, travel. This is a critical trust and licensing point for regulated customers: none of their records leak into the weights.
  • A simulation-to-reward pipeline is the core contribution. Workflow specifications are expanded into persona-conditioned, multi-turn tasks, with task-resolution rewards grounded in successful tool use for data-dependent requests. For enterprise domains those specifications are written in Agent Script, Salesforce’s declarative language for building Agentforce agents; for public tool-use domains the workflow structure is synthesized directly. The same simulation and grounded-reward machinery drives GRPO across both.

On top of RL, the team applied supervised fine-tuning using NVIDIA NeMo RL, NeMo Gym, and NeMo AutoModel — NVIDIA’s open post-training stack doing real work in a production enterprise pipeline, not a demo.

The numbers: three times fewer errors on CRM tasks

In Salesforce’s own CRM Bench — a suite of real-world tasks like updating an opportunity, routing a case, or scheduling a follow-up — Koa matches or exceeds leading model performance on CRM actions with three times fewer errors. Against the open-weight base model, the arXiv evaluation shows improvements across public tool-use, agentic-reasoning, and enterprise CRM benchmarks, with the clearest gains on multi-turn tool use. The paper is candid about placement: Koa surpasses a strong proprietary baseline while remaining below the strongest frontier models — a deliberate trade of generalist ceiling for enterprise reliability.

That positioning matters. A model that routes a support case correctly every time, with tool calls grounded in real schema, is worth more to a CIO than one that wins math olympiads.

Why control of the weights is the real story

Salesforce controls the model weights and performs post-training and inference entirely within its own trust boundary. For enterprise buyers, that converts “specialized model” from marketing language into an architecture: no customer data crosses out of Salesforce’s infrastructure during inference, and the model can be governed, audited, and versioned like any other internal asset.

The partnership extends the same principle to government. NVIDIA open models and accelerated computing are being brought into Missionforce, Salesforce’s platform for the public sector, including Missionforce Operations — a product that digitizes procurement, supplier management, and logistics workflows. Post-trained Nemotron models will run on private clouds, classified networks, and fully air-gapped systems, trained on an organization’s own operational data and terminology. Missionforce Operations is generally available now in U.S. regions; post-trained NVIDIA models for it arrive for select customers in October 2026.

From pilots to production

Koa is already running inside Salesforce — including a Slack agent that helps employees find information and complete everyday tasks — and is now moving into customer pilots with 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine, and Xero. Availability opens to select pilot customers now in Agentforce, with general availability expected winter 2026 in U.S. regions.

The pilot roster reads like a cross-section of where specialized reasoning pays off fastest: tax accounting (1-800Accountant), credit unions navigating member goals (BCU), business travel’s endless moving parts (Engine), and hospital back-office coordination (UChicago Medicine). Engine’s CEO Elia Wallen put the requirement succinctly: “what we need from AI isn’t a model that sounds confident — it’s one that can reason precisely through complex, multi-step problems.”

The bigger shift: open weights as enterprise leverage

The most consequential fact about Koa is its foundation. NVIDIA’s decision to release Nemotron as an open-weight family created a new build path for companies that have proprietary domain data but no desire to pretrain a frontier model: take a strong open base, apply specification-driven RL, and own the result outright. Salesforce is the largest company yet to walk that path in public, with a technical report and named pilot customers attached.

It also deepens NVIDIA’s strategic position beyond silicon. Every Nemotron-derived enterprise model ties NVIDIA’s software stack (NeMo, and by extension its GPUs) into customer workloads — selling shovels and the shovel-sharpening shop.

For the AI industry, Koa marks the moment “enterprise AI” stopped meaning “GPT with a business subscription” and started meaning owned, specialized, verifiably-scoped models living inside corporate trust boundaries. Expect every large software vendor watching Dreamforce this week to ask their model team the same uncomfortable question: why don’t we own ours?