No Strings Attached: Ex-OpenAI Researcher's TypeSafe Launches Jev, a Model That Refuses to Write Text
TypeSafe AI emerged from two years in stealth with Jev, the first 'System One Model' — a frontier-class decision engine that outputs typed, calibrated probabilities instead of text, at 70-500ms latency and $0.042 per million input tokens.
Every frontier lab is racing to make models that write better prose, better code, better reasoning chains. Diogo Almeida spent two years in stealth building a model that refuses to write anything at all — and on September 15, 2026, his company TypeSafe AI pulled back the curtain on the result: Jev, the first of a new class the company calls “System One Models.”
The founding question, as Almeida tells it, is one the industry keeps asking itself: models have been superhuman at chat for years, so where is all the automation? Almeida is not a casual observer of that gap. He co-authored OpenAI’s 2022 InstructGPT paper — the human-feedback training work that turned GPT-3 into a useful instruction-follower — and his name appears in the credits of the original ChatGPT release. The very techniques he helped invent are, in his telling, the reason chat models still can’t be trusted inside software: models trained to produce responses humans prefer remain overconfident, inconsistent, and occasionally unhinged, which is fine for a chatbot and disqualifying for a function call.
What a System One Model actually is
The name is a nod to Daniel Kahneman’s Thinking, Fast and Slow. Where LLMs are System 2 — slow, deliberative, chain-of-thought reasoning — TypeSafe’s models aim to be a fast, intuitive judgment layer that software can call directly. The interface is deliberately radical: developers send Jev a block of unstructured state (text, documents, program state) plus a set of typed questions, and Jev returns structured choices, scores, and probability distributions. No prose. No code. No free-form strings. TypeSafe describes it as “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.”
Three engineering choices distinguish the stack:
- A new architecture optimized for structured outputs. The space of possible outputs is defined in advance by a schema, which means the model cannot produce a type error — TypeSafe calls this mathematically guaranteed, and challenges anyone to falsify it with a single counter-example.
- A parallel sampler. LLMs generate tokens sequentially, each conditioned on the last. Jev evaluates all questions against the input state in a single pass, which is where much of its speed comes from. Adding more questions to a query barely moves response time.
- RLCD — Reinforcement Learning for Calibrated Decisions. Where RLHF optimizes for human preference and RLVR optimizes for verifiable rewards, RLCD trains the model toward calibrated answers: every output ships with probabilities and confidence scores, and higher stated confidence correlates with higher actual accuracy.
The numbers
The headline figures, if they hold up outside TypeSafe’s own testing, are striking:
- Latency: 70-500ms end-to-end, which TypeSafe frames as 40-200x faster than frontier models on decision-shaped tasks (frontier LLMs range from 3 to 329 seconds).
- Cost: $0.042 per million input tokens, versus $0.20-$10 for frontier models. Output tokens are free — “too cheap to meter.”
- Type errors: zero, by schema enforcement rather than by measurement.
- Homepage claims: up to 193.6x faster and 444.6x cheaper on one internal workflow, which the company itself says sits at the high end of real-world gains.
To demonstrate the speed, TypeSafe showed Jev playing Doom — a bot making roughly 10 model calls per second against a structured text representation of game state, at a cost of about $7 per hour. A second demo, “wikiracing,” has the model navigate between Wikipedia pages by choosing among thousands of links, a task where hallucinated link choices are fatal and where Jev’s support for up to 255-way choices per query matters.
The asterisks — and TypeSafe knows it
“We love skeptics, and are skeptics ourselves,” the launch post reads, and to its credit, the company pre-empts most of the obvious objections. The workflow evaluations backing the biggest claims are internal, built by TypeSafe’s own capabilities team, with reference answers derived from the average of GPT-6 Astra and Claude Fable 5.1 rather than independent ground truth. No paper, no weights, and no third-party evaluation accompanied the launch — RuntimeWire’s coverage notes the release describes the architecture “at a high level” only.
The “zero hallucination” claim also deserves its asterisk. Jev guarantees its output matches the requested schema — no invented fields, no malformed tool calls, no invalid types. It can still choose the wrong answer. What the calibrated confidence fields buy you is the ability to set thresholds: act autonomously above 95% confidence, route to a human below it. That’s a narrower promise than the marketing implies, but for the specific failure mode that poisons agentic systems — hallucinated calls buried deep in a dependency chain — it’s precisely the guarantee automation needs.
Access is also gated: Jev is available in early access, with developers being pulled off a waitlist rather than offered open sign-up.
Why this matters beyond one startup
The strategic bet is embedded in the name. Jev is named after William Stanley Jevons, of Jevons paradox fame — the 19th-century observation that efficiency gains increase total consumption. Almeida’s wager is that a 10-100x drop in the cost of intelligence-per-decision will cause developers to embed model calls in places where LLM economics made no sense: real-time UX paths, map-reduce over petabytes, per-event scoring, guardrails judging other models’ outputs. The model isn’t trying to be a better assistant; it’s trying to be a cheap, reliable software primitive — a “smart if-statement.”
The company itself frames the classic LLM use cases — chatbots, copilots, verifiable problems, demos — as not what Jev is for. Its lane is the fuzzy middle: classify, route, score, extract, branch — the thousands of small judgment calls currently implemented as brittle hand-written rules or as expensive, slow, occasionally-wrong LLM calls with a JSON parser bolted on.
There’s a competitive backdrop here too. This week’s OpenRouter data showed developer spend flipping back toward OpenAI as coders route Claude Code harnesses to GPT-5.6 Sol — evidence that developers treat frontier models as interchangeable commodities behind APIs. TypeSafe is betting the opposite: that the next layer of value isn’t a better general model, but a narrower one whose interface software can finally depend on. If Jev’s production numbers land anywhere near its internal evals, every platform building agent infrastructure now has a new question to answer: why is your decision layer made of strings?
Almeida is joined by co-founders Sasha Sheng, a former Meta and FAIR research engineer, and Erik Gafni, previously of genomics AI startup Ravel, Invitae, and Freenome. The company hasn’t disclosed funding. Its proof will come the way all such proofs come: not from launch-day benchmarks, but from production workloads where a 500ms, half-a-cent decision quietly outperforms a 30-second, two-dollar completion.
For now, the most honest summary is the one TypeSafe offers itself: extraordinary claims, partial receipts, and a waitlist that just opened.