← All posts / Models

16,379 Probes Later: Independent Audit Strips Jev of Its Frontier Badge

A black-box audit of TypeSafe AI's decision model Jev ran 16,379 live benchmark requests plus 3,331 follow-up probes and concluded it is not frontier at all — a small 4–9B active-parameter scorer whose 97.9% ARC-Challenge score comes from a race that finished years ago.

16,379 Probes Later: Independent Audit Strips Jev of Its Frontier Badge

Jev, the “System One” decision model from TypeSafe AI, has been marketed since its September 15 launch with a bold pitch: a frontier-class reasoner that cannot hallucinate, incredibly fast, incredibly cheap, built by the co-inventor of ChatGPT. An independent black-box audit published this week ran the model through 16,379 live benchmark requests plus 3,331 follow-up probes, and its verdict lands hard on one of those clauses: Jev is not a frontier model by any late-2026 standard. It is something smaller and humbler — and, the auditors argue, something genuinely useful in a way the marketing obscured.

What the audit did

The Jev Research team never saw weights, gradients, or serving internals. Everything was measured through the public API: answers, probability vectors, token counts, latency headers, and billing usage. The battery was substantial — 12,032 MMLU-Pro questions, 1,172 ARC-Challenge items, 196 GPQA Diamond questions, 494 multiple-choice items from Humanity’s Last Exam, the encodable MATH-500 subsets, and 167 grids from ARC-AGI-2, re-encoded two ways.

That scale matters. Small probes had already circulated — most notably Archer Hume’s “Jev’s Architecture Unmasked,” a ~28-page analysis from more than ten thousand API calls published in late September. The new audit positions itself as a short counterpoint to that work, but it goes further in several directions: it probes the probability lattice, measures tokenizer behavior script by script, runs a six-language inference test, and — importantly — estimates size.

The verdict on each claim

“Cannot hallucinate.” Technically defensible, narrowly true, misleading as sold. Jev cannot return free text — it answers only by selecting among caller-supplied options, so there is no prose to fabricate. But a contract-valid answer is not a true answer. The audit found Jev frequently wrong, and every wrong answer arrives beautifully formatted. A model that cannot write prose has not solved hallucination; it has just changed its shape.

“Fast and cheap.” Both true — and both explained by architecture rather than frontier economics. The leading theory from the timing data: prefill-dominated scoring without autoregressive decoding. Every question and option is scored in a single forward pass, at a measured floor of 73 ms plus 6.1 ms per 1,000 tokens, linear to 29k tokens. That shape of task is inherently fast and cheap; the speed is a property of the design, not proof of frontier engineering.

“Frontier-level.” This one simply does not survive. On one-shot knowledge multiple-choice, Jev lands in the band of 2025-era small instruct models: 82.7% on MMLU-Pro and 76.5% on GPQA Diamond — bracketed by Qwen 3.5 9B (82.5/77.6) and Claude 3.7 Sonnet without thinking (80.7/76.8). On anything multi-step it falls off a cliff: 21.9% on HLE’s multiple-choice subset and zero exact grids on ARC-AGI-2. The headline 97.9% on ARC-Challenge is real, but it comes from a benchmark that saturated years ago — a race that finished before 2026 began. Where Jev looks frontier-like, the audit notes, it is competing in races that are already over.

“Built by the co-inventor of ChatGPT.” Diogo Almeida was one of eight primary authors on InstructGPT and one of 88 people thanked in ChatGPT’s initial release. Deserved credit, the audit says — but “I co-invented ChatGPT” overassigns a legacy built by thousands.

How big is Jev, really?

The size estimate is the audit’s most headline-grabbing result. A slope alone is not a size — the server batches everyone’s tokens together — but two independent angles converge on a band: from the throughput side, a marginal prefill rate of ~165,207 tokens/s under compute-bound batched prefill bounds active parameters from above; from the capability side, its 2025-era-small-model profile bounds them from both directions. The result: roughly 4–9 billion active parameters, with quantization likely. Not a frontier behemoth — a small transformer with a probability read-out where a language head usually goes.

Forensics of a probability lattice

The deepest section of the report dissects Jev’s returned probabilities, and it reads like a forensic accounting exercise. Across 704,277 values in 7,887 published vectors — from 2-option questions up to 255-option menus — not one value ever landed off a 0.01 grid. The rounding has structure: sums are always 0.990 or 1.000, never above, at any option count. Independent per-value rounding would scatter sums by ±0.04 at 255 options, so a bounded correction must run after rounding, and it only ever takes mass away. Nine published vectors return a choice that is not the argmax of their own displayed table — each gap exactly one quantum — consistent with a higher-precision decision combined with separate display processing.

Option ordering matters more than buyers might expect. On a flat creative task with 254 candidate continuations, eleven orderings of the identical option set produced ten different winners, with rank agreement falling to Spearman 0.26–0.44. The practical fix the auditors use: ensemble six or more cyclic rotations and average, which restores rank agreement to 0.91–1.00 at the cost of a flatter aggregate.

The tokenizer fingerprint no one could match

Two findings deserve wider attention. First, Jev’s tokenizer is Latin-centric with byte-level fallback — roughly one token per non-Latin character across Cyrillic, Greek, Arabic, Hebrew, Thai, Devanagari, Hangul, Kana and common CJK, and for rare scripts up to ~1 token per UTF-8 byte, making number-heavy or CJK-heavy requests several times more expensive than the same volume of English. In a matched six-language XNLI inference test, Jev still outperformed Qwen3.5 9B in all six languages — 87.5% vs 83.3% in English — so the penalty is real but not disqualifying.

Second, when early probes let Jev describe itself, it identified itself with an OpenAI-family name 337 times out of 360 — but the auditors read this as a learned assistant-style prior from internet text, not a signature of authorship. It accepts a fictional origin story at 0.97 if you assert one in the prompt. The model knows nothing about who it is; its self-narrative is inherited from training data rather than lived identity.

Cheap enough to rewire how you call models

The Pareto analysis is where the audit’s contrarian warmth lives. Jev is not frontier — but due to the cost savings from a lack of text decoding, it does move the frontier of price-performance, “and then some.” The entire 12,032-question MMLU-Pro run cost $0.28 — about $0.000024 per question, four to five orders of magnitude cheaper than vals.ai-measured flagships. The same run at Claude Fable 5.1’s measured per-test cost would run about $1,900+.

That price point defines the product shape the auditors think Jev actually occupies: routing a menu of a hundred intents, grading answers against a rubric, scoring which of twenty rewrites a user most likely meant, gating a pipeline on document relevance. Fast, deterministic-shaped, cheap structured judgement — not conversation, not research.

What TypeSafe should take from this

The audit does not call Jev a fraud — it calls the marketing a category error. The model earns attention by being useful, not by being frontier. The report’s practical guidance for builders: treat the choice as the answer and the vector as a ranking signal (calibration is dataset-dependent — near-calibrated on MMLU-Pro, badly overconfident on HLE); average repeats where stability matters, since identical calls differ; normalize option surface forms, because a leading space can change the winner; and pack questions into shared requests, where 192 questions cost ~153 ms versus ~280 ms each sent separately.

For an industry mid-conversation about what “frontier” even means — and a week in which the White House accord institutionalized self-policing of frontier development — a 16,379-probe audit demonstrating that a heavily-marketed “frontier” model is a small, well-made decision engine is a useful corrective. Jev’s real competition isn’t Claude or Gemini. It’s the if-statement.