← All posts / Models

38 Milliseconds to Decide: Cloudflare Open-Sources Clef, Its First Homegrown AI Models

Cloudflare's first self-trained models, Clef and Clef-flash, are Jev-compatible decision models with vision, a 64K context window, and median latency as low as 38.8 ms — released under Apache 2.0 with a new RL fine-tuning service.

38 Milliseconds to Decide: Cloudflare Open-Sources Clef, Its First Homegrown AI Models

Cloudflare has spent years renting out GPUs at the edge and hosting other people’s models on Workers AI. On October 1, 2026, during its Birthday Week, the company crossed a threshold it had never crossed before: it shipped models it trained itself. Clef and Clef-flash are “decision models” — a fast-emerging class of AI that doesn’t write prose, doesn’t reason out loud, and doesn’t stream tokens. You hand them a state and a schema of typed questions, and they hand back calibrated probabilities you can act on immediately. Both are open-weight under Apache 2.0, both are live on Workers AI, and the smaller one answers in a median of 38.8 milliseconds.

The launch puts Cloudflare in direct competition with Typesafe’s Jev, the decision model that has been absorbing the industry’s attention since its debut — and Cloudflare is picking the fight on the two axes that matter most for this category: latency and openness.

What a decision model actually is

The core insight behind the category is that most of what agents do with LLMs is not generation — it’s judgment. Should this ticket be escalated? Is this request urgent? Which team owns this? Today, teams burn a frontier model on those calls, parse free-form text, tolerate non-determinism, and wait for reasoning tokens to finish streaming. A decision model inverts the contract: it reads an input state plus up to 64 typed questions, and returns a probability for every allowed answer. There is no text to parse and nothing to wait for.

Clef supports three question types. noul is a yes/no question that returns the probability that the answer is yes. choice picks one option from a set you define, returning the selection, a per-option probability, and a confidence value. score rates against an ordered rubric and returns a probability-weighted score. Crucially, Clef follows the same System One API as Jev, so an existing Jev integration can switch by changing the endpoint and model name — a deliberately frictionless migration path.

The name is a Cloudflare in-joke with real logic behind it: in music, a clef assigns meaning to the lines and spaces that follow it. A decision model does the same for an agent’s subsequent actions. And “CF,” of course, hearkens to Cloudflare.

The numbers

Against Jev, the latency story is blunt. Across 43 benchmark runs, Clef posts a median of 209.3 ms — 2.5× faster than Jev’s 524.1 ms — while the 9B Clef-flash lands at 38.8 ms median, roughly 13× faster. At the p95 tail: 238.6 ms, 122.4 ms, and 536.0 ms respectively. Because the models are hosted on Workers AI, inference runs on GPUs inside Cloudflare’s network, close to users, keeping the network round trip short. That combination is what makes it plausible to put a model decision directly in the hot path of a request — block the request, route the ticket, or escalate to a human, then hand off to a full LLM only when one is genuinely needed.

On quality, Cloudflare claims leadership of the Jev Decision Index: across 10 decision benchmarks, a Clef variant scores highest on 7. BFCL (case exact) comes in at 98.47 for Clef and 98.76 for Clef-flash against Jev’s 95.75. BANKING77 macro-F1 is 94.20 versus 79.74. Home appliances classification is a blowout — 97.73 for Clef-flash against 52.27 for Jev. The losses are honest, too: Jev retains the edge on When2Call accuracy (80.97 vs 72.37) and BRIGHT retrieval, and DiffusionGemma Jev still wins PhishNChips phishing detection at 85.35. On Typesafe’s own workflow evals, Clef beats Jev in 3 of 4 areas — invoice processing, customer service, and security incidents — while Jev keeps agent trace observability.

Two structural advantages compound the scores. Clef’s context window is 64K tokens, double Jev’s 32K, so there’s more room for the state being classified. And Clef has a vision encoder, accepting up to four images alongside text — something no text-only decision model, Jev included, can do today. Visual classification of screenshots, product photos, or UI states becomes a single-call operation.

Cloudflare’s own threat intelligence team has been running the model in production-adjacent workflows: paired with Browser Run, Clef fetched, rendered, and classified a domain in 2.2 seconds, versus 4.7 seconds for gpt-oss-120b in the identical pipeline — and returned richer, probability-ranked category labels instead of two classifications.

How it was trained

The technical backstory is one of the more interesting parts of the release. When Jev shipped, Cloudflare had already been experimenting publicly with building deterministic probability outputs on top of a diffusion language model — work that built on independent research by Matt Mastracci and on upstream vLLM contributions. Clef keeps the concept but changes the foundation: a frozen Qwen backbone, Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash.

The inference design is the key to the speed. Clef runs Qwen as a prefill-only pass, then scores all valid schema choices in parallel. The decision step is non-autoregressive — no intermediate text is generated token by token, because answers are derived directly from internal backbone representations via a two-stage attention routing process: each valid choice extracts prompt-relevant context, and individual field parameters cross-attend with other fields and back to the original payload before scoring. Training froze the backbone and jointly optimized the routing head with rank-256 low-rank adapters, using label-smoothed cross-entropy for valid schema outputs plus a Brier loss for probability calibration, over internal synthetic datasets that permuted field orders, prompts, and schema structures. A secondary optimization target — Reinforcement Learning for Calibrated Decisions (RLCD) — grants partial credit for adjacent ordinal choices, rewards fully precise records, and applies a reference penalty against distribution shift.

Alongside the models, Cloudflare is debuting an RL fine-tuning service: hands-on tuning with a forward-deployed engineer team first, self-serve capture-train-redeploy platform later. The company is explicit that its 15-plus years of network data — bot classification, support triage, trust-and-safety labels — is the moat it intends to exploit for domain-specific variants. Hosted pricing is aggressive for the category: reported at roughly $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with an enterprise promise that Cloudflare does not read, store, or train on requests.

Why it matters

The decision-model category is filling in the middle of the AI stack: frontier LLMs for reasoning and generation, tiny calibrated classifiers for the millions of judgment calls in between. What Cloudflare adds to that picture is the edge. A 38.8 ms median decision from a GPU physically near the user changes which decisions are worth automating — guardrail checks and routing calls that were too latency-sensitive for an LLM round trip now fit inside the request path.

And the Apache 2.0 release changes the competitive math for Typesafe. Jev’s momentum was real, but a drop-in-compatible alternative that is faster on most benchmarks, open enough to self-host, backed by one of the largest GPU fleets at the edge, and paired with a fine-tuning service is the kind of commoditization pressure that reshapes a young market. For builders, the immediate takeaway is simpler: the cost of adding fast, calibrated judgment to an agent just dropped again — to about a tenth of a second, and about a dime per million tokens.