← All posts / Models

One Day, One LoRA, 90.1%: Bespoke Labs Open-Sources the Entire Jev Recipe

Bespoke Labs publishes the data, model, and training code for Nimble-9B — a one-day LoRA on Qwen3.5-9B that lands 3 points behind closed Jev on its own style of eval, and the real lesson is contrastive data curation, not scale.

One Day, One LoRA, 90.1%: Bespoke Labs Open-Sources the Entire Jev Recipe

When TypeSafe AI launched Jev on September 15, 2026, it pitched a new category: the “System One Model” — no text generation, just typed decisions (a choice, a boolean, a rubric score) computed by scoring the allowed answer tokens in a single forward pass, fast and cheap enough to sit inside an agent loop. What it did not publish was the recipe. No weights, no data, no training code.

Five days later, the open-source community has published one for them. Bespoke Labs released Bespoke-Nimble-9B under Apache 2.0 — and unlike most open clones, the release is not just a checkpoint. It is the full trifecta: the curated training data, the LoRA adapter, and the exact recipe that produced it. The README states the engineering constraint plainly: “We built Nimble in one day, so expect some rough edges.” On the project’s own 324-example held-out evaluation, the model agrees with reference labels on 90.12% of cases, versus 93.21% for Jev 1.13.0 accessed through its API — a 3.09-point gap, built with a ~165 MiB adapter on a 9-billion-parameter base model.

What Nimble actually is

Nimble is a LoRA (Low-Rank Adaptation) fine-tune of Qwen3.5-9B, not a model trained from scratch. The adapter is small — about 165 MiB — and the base checkpoint is pulled separately from Hugging Face. The serving mechanism is the interesting part: following the decoding approach that researcher Niels Rogge documented in his viral post on how Jev works, Nimble prefills the KV-cache with the context and schema once, then reads the logits for the allowed answer tokens directly. There is no decoding loop, no generated JSON to parse, no chain of thought. The softmax over candidate tokens is the output, complete with per-answer probabilities.

The contract is deliberately narrow. A schema must be flat — no nested fields — with each field being either an enum (1 to 26 string choices) or a boolean. Prompts over 2,048 tokens are rejected rather than truncated. It accepts text only, even though the base model has a vision component, and it cannot write explanations or extract spans from context. In exchange, it runs anywhere: an NVIDIA GPU with PyTorch, or a MacBook with Apple Silicon via MLX, where a ParallelScorer processes the shared context once and scores all fields in parallel. There is also a hosted option on Modal and a public API for anyone who wants to try before downloading.

On latency, the project’s own numbers: a median of 106 ms per example on an H100, and a median of 444 ms on an M5 Pro MacBook with 64 GB of memory — no server required. For comparison, the same table lists Jev’s API at a 246.7 ms median, and general-purpose frontier models generating text the usual way (via OpenRouter) at roughly 2.8 to 15 seconds per example.

The numbers, honestly framed

On 324 held-out examples — 162 contrastive pairs drawn from six source families — the comparison looks like this:

ModelAgreement
Gemma 3 270M IT28.70%
Qwen3.5-0.8B45.37%
Qwen3.5-4B61.42%
Qwen3.5-9B (base)66.36%
Qwen3.8-27B (untuned)84.88%
Bespoke-Nimble-9B90.12%
Jev 1.13.0 (API)93.21%

Two readings of this table matter more than the headline. First, the fine-tuned 9B model beats the untuned Qwen3.8-27B by 5.25 percentage points — a model three times its size. Targeted training beat a tripling of parameters. Second, Jev’s remaining lead is 3.09 points, and Jev is a from-scratch system with an undisclosed training pipeline behind it. Whether that gap survives on distributions outside this eval is exactly the question the repo cannot answer by itself — the held-out set is narrow, drawn from just six source families, and every reference label is synthetic. The authors are unusually candid about this: “All of the labels are synthetic: a model checked them, and no person has reviewed them.” They also point to a separate guide that runs Nimble and Jev side by side on thirteen public, human-labeled benchmark subsets, starting with VitaminC.

Contrastive data curation: the part worth stealing

The recipe’s core is a technique Bespoke calls contrastive data curation, and it is the mechanism behind the jump from the base model’s 66% to 90%. The idea: build training examples in minimal pairs that differ in exactly one relevant fact — a change that flips the correct answer while everything else, including the question and the policy, stays identical. In the repo’s canonical example, a refund is authorized only if the sole authorization was signed by someone with approval rights. In one example the signature is Mira’s (who has rights): answer true. Change one name to Noah: answer false. The model learns precisely which evidence should change a decision, and — because negative examples are constructed rather than merely collected — it is pushed toward calibration, not just accuracy.

The pipeline has four steps, each with teeth. Check the decision rules against the schema. Build the pair, changing at most eight words. Then verify both examples with separate model calls — and crucially, delete each evidence sentence in turn to confirm the focus fact becomes unknowable without it, so no other text leaks the answer. Finally, generate labels by code applying checked rules to checked facts, and keep a pair only when every check passes and the two labels differ. Every model request and response is saved so the whole process can be replayed offline.

The total scale of this effort is almost comically small by frontier standards: 2,676 training examples across ten subject categories, one epoch of LoRA at rank 16, learning rate 5e-5, effective batch size 8 — tuned on an L40S, with the final fit and evaluation on an H100. The authors note they explicitly did not distill from Jev; hard reference labels were generated by their own pipeline, with saved Jev probabilities available only as an optional future soft-target experiment.

The 48-hour replication wave, contextualized

Nimble is one of at least six independent open implementations that appeared within roughly 48 hours of Jev’s launch — a wave catalogued by Latent.Space that also includes Laya (a 421M-parameter PPO-trained model from a researcher who had published the same core idea on arXiv in March 2025), Jevlike (a ~40 KB embedding-only implementation), Kev-0.5B (built to run offline on a MacBook), OpenJev, and DiffusionGemmaJev. Resource levels span four orders of magnitude, yet several claim to approach or match the closed original on its own terms.

The signal worth tracking is not any single score — every number above is self-reported — but the replication speed. A frontier lab’s novel-sounding capability was decomposed, reimplemented, and published with full provenance inside two days, at least for the narrow output shape Jev targets. The bottleneck was never architecture; it was knowing that the trick works, after which a competent team with an existing base model and a data-curation discipline could close most of the gap. For builders, the practical takeaway is straightforward: before paying API premiums for a closed decision model, price out a LoRA on an open base with contrastively curated pairs for your own task distribution. The economics may surprise you — and if it doesn’t close the gap, the closed model will still be there, now with a real baseline to beat.

Model specs, scores, and timelines reflect the repositories and reporting as of September 20, 2026; all evaluation numbers are the project’s own self-reported results, not independently verified.