← All posts / Tools

One Loop to Ship Them All: CoreWeave Forge Turns Production Runs Into Better Models

CoreWeave's new Forge development layer unifies run, observe, curate, improve and evaluate into one open environment, with a coding agent named ARIA and serverless RL that trains 1.4x faster at 40% lower cost.

One Loop to Ship Them All: CoreWeave Forge Turns Production Runs Into Better Models

The most expensive waste product in modern AI is not GPU hours — it is signal. Every day, production agents generate millions of traces that reveal exactly where a system fails, and every day those traces evaporate into observability dashboards that no training pipeline ever reads. CoreWeave thinks it has found a way to stop throwing that signal away, and it has built an entire product layer around the idea.

At its Fully Connected conference in San Francisco on September 30, 2026, CoreWeave (Nasdaq: CRWV) announced CoreWeave Forge, a development layer that runs the entire AI improvement loop — run, observe, curate, improve, evaluate, and repeat — in one connected environment. The pitch is deceptively simple: what a business learns from running AI in production should flow directly back into how engineers make the next version better. MasterClass and Canva are already building on it.

The fragmentation problem

Model and agent development in 2026 is a relay race run with a dropped baton at every exchange. Training runs live in one vendor’s experiment tracker. Evaluations live in another. Agent traces live in a third observability stack. None of them were built to talk to each other, so a production trace never feeds the next training run, a finished experiment never informs the next evaluation, and every handoff between tools loses signal or burns engineer time.

Forge is CoreWeave’s answer: unify Weights & Biases Models, post-training expertise from OpenPipe, and the open-source marimo notebook project with CoreWeave’s own services into a single environment built for continuous improvement. All three of those are pieces CoreWeave acquired or absorbed earlier in 2026, and Forge is the moment the acquisitions snap together into a coherent product rather than a portfolio.

Notably, CoreWeave is selling openness as loudly as integration. Forge stays open across the models, frameworks, and other clouds a team already uses — a direct jab at hyperscalers whose toolchains quietly assume you will stay inside their walls. “Every platform decision in this market has involved weighing how much optionality a team is willing to surrender for a connected toolchain,” said Nick Patience, vice president and practice lead for AI platforms at the Futurum Group. “CoreWeave Forge is built on the premise that teams shouldn’t have to make that calculation.”

What actually ships

Forge arrives with a deep stack of new and generally available capabilities spanning the loop:

  • CoreWeave ARIA (GA) — a coding agent that automatically analyzes large-scale experiment data and agent observability data, surfaces what drove a change, and proposes the next experiments worth running, with the evidence to back it up.
  • Weights & Biases Models — tracks and compares tens of thousands of experiments and millions of metrics, with advanced visualizations and automated workflows.
  • CoreWeave Notebooks — a developer experience built on marimo that lives where the work runs, so a prototype carries straight into training, evaluation, and production “instead of dying in a rewrite.”
  • Agent Lens — intelligent observability for production agents that analyzes tens of millions of traces and turns them into proven fixes. CoreWeave claims it detects 20 percent more critical failures and fixes issues at one-tenth the cost compared against a general-purpose frontier LLM on the same benchmark.
  • Sandboxes (GA) — a fresh, isolated environment for every run, whether agent tool use, reinforcement learning, or evaluation, pre-connected to the rest of the loop.
  • Registry — every model checkpoint and agent configuration versioned in open, portable formats with clear lineage, so the next attempt can start from any point in a team’s history, on any stack.
  • Post-Training — serverless RL, serverless supervised fine-tuning, and model distillation. CoreWeave says its serverless RL trains 1.4 times faster at 40 percent lower cost than a self-managed setup.
  • Inference — playground and deployment access to the latest open-weight models, plus a new preview capability called RL Rollouts that hot-loads checkpoints into a live deployment so a training loop can keep running without redeploying.

The RL Rollouts detail deserves a pause. Hot-swapping checkpoints into live inference is one of those unglamorous infrastructure tricks that quietly changes how teams work: instead of the traditional train-freeze-redeploy cycle that kills iteration momentum, improvement becomes something closer to continuous. Combined with Registry’s versioned lineage, a team can A/B a newly fine-tuned checkpoint against the live one without standing up parallel infrastructure.

Why a cloud company is building dev tools

The strategic logic is worth spelling out. CoreWeave’s core business is renting AI compute, and that business is commoditizing from two directions at once — hyperscalers on one side and a proliferation of neoclouds on the other. Raw GPU capacity no longer differentiates; the software layer above it increasingly does. CoreWeave has leaned on this story all year, pointing to record MLPerf inference and training results and its status as the only AI cloud to earn the top Platinum ranking in SemiAnalysis ClusterMAX three times in a row.

Forge extends that logic from infrastructure to workflow. If the improvement loop lives on CoreWeave — experiments, traces, checkpoints, evaluations, and the post-training jobs that consume them all in one place — then the compute those jobs run on is remarkably sticky. A team that curates production traces into training data inside Forge has a rising cost of leaving, not because of lock-in file formats (Registry is explicitly open and portable) but because the workflow itself compounds value over time.

There is also a market-timing argument. “As more people build with AI, they’re putting models to work with their own data and workflows,” said Chen Goldberg, executive vice president of product and engineering at CoreWeave. “That’s where the gaps between model capability and system performance become clear.” The frontier-model race gets the headlines, but the bulk of 2026’s AI spend is flowing into exactly the layer Forge targets: teams adapting open-weight and frontier models to proprietary data, and struggling to measure whether version N+1 is actually better than version N under real operating conditions.

Fine print and open questions

Forge ships in three editions — Free, Pro, and Enterprise — with a 30-day free trial of Pro, so teams can “start without a procurement cycle.” The Free tier is aimed at personal development; pricing beyond that scales with usage of the underlying loop.

The claims worth scrutinizing are the benchmark ones: 20 percent more critical failures detected and one-tenth the fix cost for Agent Lens, and 1.4x faster RL training at 40 percent lower cost. All are self-reported against baselines CoreWeave chose (a “general-purpose frontier LLM” and a “self-managed setup” respectively), and the company has not yet published the benchmark methodology in full. Buyers should demand reproduction before accepting them as planning inputs.

The competitive response will be the thing to watch. Weights & Biases was the independent experiment-tracking standard before CoreWeave absorbed it; its rivals — and the observability vendors whose lunch Agent Lens is explicitly eating — now have a former neutral platform owned by a cloud competitor. Expect the pitch “your tool vendor shouldn’t also be your cloud vendor” to appear in a competitor’s deck near you.

What is not in question is the direction. The era of treating model development and model operations as separate disciplines is ending, and the companies that win the next phase of AI infrastructure will be the ones that close the loop between what a model does in production and what its successor learns from. Forge is one of the most complete attempts yet to build that loop as a product — and it signals that CoreWeave intends to compete on software, not just on the size of its clusters.