← All posts / Tools

TrueFoundry Open-Sources TrueForge: A Vendor-Neutral Agent Harness That Undercuts Claude Managed Agents by Up to 75%

MIT-licensed agent harness matches Claude Managed Agents' accuracy at 30-75% lower cost, and a $10,000 hackathon kicks off today.

TrueFoundry Open-Sources TrueForge: A Vendor-Neutral Agent Harness That Undercuts Claude Managed Agents by Up to 75%

The agent layer of the AI stack is consolidating fast, and TrueFoundry just made the most aggressive move yet to keep it open. On August 19, 2026, the enterprise AI infrastructure company released TrueForge, an MIT-licensed, open-source agent harness that runs production AI agents on your own infrastructure with any model you choose. Five days later, the release has become the center of gravity for a developer community event — The Agent Harness Hackathon, run by WeMakeDevs, which kicks off today (August 24) and runs through August 30.

What TrueForge Actually Is

Strip away the marketing and most “agent platforms” resolve to the same anatomy: a model, a set of tools, and a harness that orchestrates the loop between them. The harness decides how context is assembled, when tools get called, how subagents are spawned, how code execution is sandboxed, and how human approvals are injected. Anthropic’s Claude Managed Agents bundles that loop as a paid, closed service. TrueForge is TrueFoundry’s answer: run the same kind of loop yourself, on your own infrastructure, with the model of your choice — including open-weight models.

According to the GitHub repository, TrueForge handles the full operational loop: model calls, MCP (Model Context Protocol) tool integrations, skills, sandboxing, approvals, and context management. It is explicitly vendor-neutral: there is no lock-in to a single model provider, and the harness itself is inspectable, forkable code rather than a black-box API. The project pairs with TrueFoundry’s commercial AI Gateway, which handles enterprise concerns like routing, spend governance, and compliance — but the harness works standalone.

The Benchmark Claims: Same Accuracy, 30-75% Cheaper

The number that got attention is cost. TrueFoundry published a head-to-head benchmark on what it calls Enterprise-Bench: 14 production-style agent tasks, run end to end. The headline result is that TrueForge matched the mean task completion rate of Claude Managed Agents running Opus 4.8, at roughly 30% lower list-price cost — $8.50 per run versus $11.80.

The more aggressive 75% figure comes from switching the model underneath. When TrueFoundry ran the same harness with a cheaper model (a DeepSeek-class model, according to third-party coverage), the cost per run dropped to around $2.90 — roughly 75% below the Claude Managed Agents baseline at a similar solve rate. Hence the “30-75% cheaper” range in the launch coverage: 30% if you keep Anthropic’s flagship model, up to 75% if you exploit the vendor-neutrality and swap in a lower-cost model.

The cost delta is not purely model pricing. The benchmark data also shows TrueForge is simply more economical in how it drives the loop: coverage reports indicate it averaged about 19 tool calls per task versus 32 for Claude Managed Agents, and consumed roughly 30% fewer tokens overall at equal accuracy. Fewer, better-aimed tool calls mean less context stuffing, fewer round trips, and a smaller bill at the end of the month.

Independent reviewers have started poking at the claims. A review from Wavect noted the published 14-task benchmark reports “equal mean task completion to Claude Managed Agents on Opus 4.8 at about 27% lower list-price cost” — close to TrueFoundry’s own numbers, though measured at list prices rather than negotiated enterprise rates. Community discussion on Reddit has surfaced similar tool-usage deltas from users running their own comparisons.

Why the Harness, Not the Model, Is the New Battleground

The strategic significance goes beyond one product launch. As frontier model prices fall — a trend we’ve covered as the frontier price war — the marginal cost of an agent run increasingly comes from the harness layer: how many tokens get stuffed into context, how many tool calls get made, how efficiently subagents are delegated to. TrueFoundry’s argument, articulated in its launch essay “The Agent Harness Should Be Open,” is that this layer belongs to the customer, not the vendor.

There is a direct competitive logic here. Claude Managed Agents bills customers for Claude model usage plus $0.08 per agent runtime hour — a meter that runs whether or not your agent is being smart about its tool calls. An open harness turns that meter into something you can profile and optimize. And with the benchmark showing equal accuracy on production-style tasks, the question enterprises are asking shifts from “which model?” to “who controls the loop?”

It also fits the broader pattern of 2026: Cloudflare’s agentic web stack, the MCP ecosystem’s rapid standardization, and the steady flow of open-weight models from Alibaba, DeepSeek, and others. Every one of those trends makes a vendor-neutral harness more valuable, because every one of them increases the number of models and tools you might want to wire together.

The Hackathon Kicks Off Today

Timing matters, and TrueFoundry’s community play is deliberate. The Agent Harness Hackathon, organized by WeMakeDevs, runs August 24-30, 2026 — online globally, with a live day in San Francisco. The prize pool is $10,000, headlined by an NVIDIA DGX Spark personal AI supercomputer (worth roughly $5,000) for the best use of TrueForge, plus a Mac Mini and job opportunities. Participants build agents that use real tools, safely execute code, and delegate tasks across subagents — exactly the skills the harness is designed to teach.

For developers, it’s a low-stakes on-ramp to a piece of infrastructure that may become table stakes. For TrueFoundry, it’s a funnel: hackathon participants become contributors, and contributors become the community that keeps an MIT-licensed project alive after the launch buzz fades.

Caveats Worth Noting

Skepticism is warranted on three fronts. First, the benchmark is vendor-run. Fourteen tasks is a meaningful sample for a demo, but it is not an independent eval, and “Enterprise-Bench” is TrueFoundry’s own construction. Independent replications exist but are early. Second, “up to 75% cheaper” requires swapping away from Anthropic’s model — accurate, but the framing compresses two different comparisons into one number. Third, running your own harness means owning your own reliability: sandboxing, approval workflows, and context management are now your operational burden, not your vendor’s. The $0.08/hour fee Claude charges is partly an insurance premium.

Still, the direction is clear. Open harnesses with published benchmarks put pressure on closed agent platforms to justify their pricing, and give enterprises genuine leverage in negotiations they didn’t have a year ago. TrueForge is on GitHub now, MIT-licensed, waiting to be stress-tested — and this week, several thousand hackathon participants will be doing exactly that.