← All posts / Industry

OpenAI-Backed Harvey Builds Its First In-House Model on China's Kimi K3

Legal AI leader Harvey post-trained Harvey Tenet on Moonshot AI's open-weight Kimi K3 — the clearest sign yet that Chinese open models are becoming the foundation layer for Western enterprise AI.

OpenAI-Backed Harvey Builds Its First In-House Model on China's Kimi K3

OpenAI-Backed Harvey Builds Its First In-House Model on China’s Kimi K3

When Harvey — the San Francisco legal-AI company backed by OpenAI’s startup fund and valued at $11 billion as of March — decided to build its first in-house model, it faced a foundational choice: which foundation? The answer, disclosed on August 20, 2026, surprised much of the industry. Harvey Tenet, the company’s first post-trained open-weight model, is built not on GPT, not on Claude, but on Kimi K3 — the 2.8-trillion-parameter open-weight model from Beijing’s Moonshot AI.

For two years, the Western AI narrative has been that Chinese open-weight models are cheap imitators: fine for hobbyists, benchmarks, and side projects, but not something a regulated Western enterprise would stake its core product on. Harvey’s move quietly dismantles that assumption. One of the fastest-growing enterprise AI companies in the world, serving a majority of top US law firms, looked at the menu of foundation models and picked a Chinese one — then trained a legal specialist on top of it that, by Harvey’s own benchmarks, beats US frontier models on long-horizon legal work at a fraction of the cost.

What Harvey actually built

Harvey Tenet is a research preview, not a shipped product — but the technical details in Harvey’s engineering post are substantial.

The base is Kimi K3, Moonshot’s flagship mixture-of-experts model released in July with a 1M-token context window, 896 experts (16 active per token), and downloadable weights. On top of that, Harvey post-trained via asynchronous reinforcement learning in realistic legal work environments, working with Fireworks’ research team. The training corpus combined synthetic data, publicly available legal data, and human expert data — Harvey collaborated with Mercor to scale the expert datasets, and emphasized that no customer data was used.

The training environments are worth pausing on, because they explain the results. Each task is a closed-universe “client matter”: a short instruction written as a request from a partner (averaging ~50 words), a matter file containing key and peripheral documents, and an expert rubric decomposing partner-level review into atomic binary criteria — facts, conclusions, citations, severity ratings. A single training rollout can span more than 1,000 turns and hundreds of thousands of tokens. The agent is initialized in a sandboxed workspace, must search the matter, read documents, and draft work product to disk, and is graded by LLM-as-judge over the rubric (Harvey converged on Kimi 2.6 as the optimal judge). Rewards were shaped not just for quality but for token efficiency — preferring trajectories that consume fewer inference tokens at equal performance.

The scale is serious for a “research preview”: GSPO with a rank-64 LoRA over the full Kimi K3 network, ~1,750 agentic legal task environments, ~150 NVIDIA B300 GPUs running for two months, with weights hot-loaded into rollout deployments so generation never stops.

The numbers

On Harvey’s open-source Legal Agent Benchmark (LAB) — over 1,200 tasks across 24 practice areas — Tenet completes almost twice as many held-out tasks as base Kimi K3, and 20% more on LAB Contracts, raising all-pass rates by 9 and 2 percentage points respectively. It achieves state-of-the-art on LAB Contracts and places second on LAB overall.

Two results stand out beyond LAB. First, the gains transfer: on Mercor’s APEX Agents (Corporate Law) and Crosby’s Redline Bench — benchmarks Tenet never saw during training — it substantially outperforms the K3 base. Second, the legal-knowledge floor held: on parametric-reasoning and knowledge benchmarks like APEX v1, Scale’s Professional Reasoning Benchmark, LegalBench, CUAD, and MAUD, Tenet maintained strong scores. Agentic training improved the lawyer-agent without eroding the law student underneath.

And then there is cost. Because open weights carry cheaper per-token prices, and because Harvey explicitly co-optimized for token consumption, Tenet pushes the quality-cost Pareto frontier outward: significant performance gains at stable cost. In companion work, a post-trained GLM-5.2 model for Harvey’s Review Table improved answer quality by 3.6 points and citation quality by 12.1 points at roughly one-tenth the cost per cell.

Three bigger threads run through this announcement.

First, the open-weight stack is going vertical. Harvey’s research didn’t stop at one model. For M&A diligence — single tasks requiring traversal of up to 80M tokens of dataroom context — Harvey extended its harness with Recursive Language Models (RLMs), where a GLM-5.2 root agent holds the dataroom in a REPL and delegates to sub-agents; post-training within that harness lifted criteria pass rates from 43.8% (best baseline) to 60.1%. For firm-knowledge search, a Qwen3.8-27B model was trained to internalize a firm’s corpus into parametric memory, cutting cost per query by 90% and tripling “intelligence per token” (190.8 vs 129.3 for the best frontier configuration). The pattern: Chinese open bases — Kimi K3, GLM-5.2, Qwen3.8 — are becoming the default substrate for Western domain specialists.

Second, sovereignty over intelligence is now a product feature. Harvey frames its research agenda as letting “law firms build their own specialized models and own their intelligence.” For a profession bound by privilege, confidentiality, and jurisdictional data rules, owning the weights — rather than renting them per token from a US lab — is a genuine selling point. It also decouples Harvey from the pricing cycles of OpenAI and Anthropic, both of which have been repricing their APIs this summer.

Third, the geopolitics are awkward and unavoidable. An OpenAI-backed company building its core research on a Beijing lab’s model is exactly the kind of story that fuels Washington’s anxieties about Chinese model adoption — US lawmakers have already grilled DoorDash over employees using Chinese models. SCMP’s headline framed it as a “pivot”; Harvey frames it as cost-efficiency. Both framings capture something true. Moonshot itself publicly congratulated the Harvey team. The uncomfortable fact for US labs: their most sophisticated enterprise customers are concluding that the best price-performance foundation for serious vertical work now comes with Chinese weights.

Caveats worth keeping

The results are research-preview grade: mostly self-reported, on Harvey’s own benchmark (LAB is open-source, which helps), with base-model scores taken from Vals’ leaderboard, and Harvey itself notes its internal runs diverge from other external reporters due to harness and judge differences. On the hard subset of LAB, improvement over base K3 was non-statistically significant (36.0% → 36.8%). Independent replication — particularly the claim of beating US frontier models on long-horizon legal work — will determine whether Tenet is a milestone or a marketing beat.

There is also a deployment question the announcement doesn’t fully answer: Kimi K3’s license reportedly includes revenue-share and non-commercial restrictions introduced this month, and running a 2.8T-parameter MoE in production is an infrastructure project in itself. Harvey says it is “focused on scaling compute” to bring this work into production — which tells you production isn’t here yet.

The takeaway

Strip away the benchmark tables and one decision remains: when the leader in legal AI needed a foundation to own and build on, it chose an open Chinese model over every closed American alternative. Whatever Tenet’s final production fate, that choice is the signal. The frontier of applied AI is increasingly being built one layer up from the frontier labs — and the layer underneath is more often than not open-weight, and increasingly, Chinese.