← All posts / Tools

One Question, 150 Milliseconds: OpenAI's Decisions API Turns Bounded Choices Into a Primitive

At DevDay 2026 OpenAI shipped the Decisions API: GPT-6 Luna focused on user-defined questions with fixed answer sets, returning calibrated decisions in ~150 ms. It formalizes the decision-only category TypeSafe's Jev created and Laya open-sourced.

One Question, 150 Milliseconds: OpenAI's Decisions API Turns Bounded Choices Into a Primitive

On a keynote stage packed with twenty-plus launches, the sleeper announcement of OpenAI’s DevDay 2026 was arguably one of the smallest: the Decisions API, a limited-preview endpoint that focuses GPT-6 Luna — the smallest, cheapest model in OpenAI’s current lineup — on “a specific set of user-defined questions with finite pre-defined answers.” No prose. No output tokens to parse. You define the answer space, send text or image context, and get back a decision with a confidence score in roughly 150 milliseconds, a round-trip that regular GPT-6 Luna takes about 1.6 seconds to complete.

Among the Dots agents, the $500 Pro tier, and Codex-in-the-cloud, a classification endpoint sounds mundane. It is not. It is OpenAI formally productizing a category that, until this month, existed only as startup products and open-source projects — and it changes the economics of the most common job in production AI systems: making a bounded choice.

What actually shipped

The Decisions API is now in limited preview, with broad rollout promised “in the coming days.” Per OpenAI’s DevDay recap and early coverage, the shape of the thing is simple:

  • You supply context (text or images), a question, and a finite set of allowed answers.
  • The service scores the options and returns one answer plus a calibrated confidence score.
  • Answers are schema-valid by construction — there is no generation step, so there is nothing to parse and no way to go off-schema.

OpenAI’s own framing: Decisions API “enables real-time decision-making by focusing Luna’s intelligence on a specific set of user-defined questions with finite pre-defined answers.” Developers can use it to classify content, route requests, or choose an agent’s next action from a limited menu.

The latency claim is the headline technical number: 150 ms versus 1.6 seconds for the same judgment through vanilla Luna. What is not published yet matters just as much: per-call pricing, the maximum number of candidate answers per request, and whether teams can tune the model on their own data. An OpenAI spokesperson told The New Stack that more details arrive “at broad rollout.” Those three unknowns will decide whether this becomes a standard building block or stays a niche tool.

Why this is a category, not a feature

The pattern the Decisions API formalizes is old: most teams today handle classification and routing with a general-purpose chat model, a carefully worded prompt, and prayer. You ask the model to pick from a list, maybe read token probabilities to approximate a confidence score, and hope the output validates. It works, it burns tokens, and the confidence numbers are a rough guess.

The alternative was training a small classifier — fast and cheap per call, but it needs labeled data and a retraining run every time the label set changes.

Decision-only models sit between the two: new labels travel in the prompt, but the score that comes back is one your code can act on. TypeSafe’s Jev popularized the architecture — parallel scoring of typed answers instead of autoregressive generation — and its sudden adoption made the category impossible to ignore. Laya then open-sourced the idea under Apache 2.0. OpenAI’s entry reads as a direct answer; multiple outlets, and the community, immediately tagged it a “Jev killer.”

The comparison table now looks like this (figures from the CMU JEV-as-a-Judge study, arXiv 2609.26550, and vendor self-reporting):

DimensionFrontier LLM (GPT-6 class)Jev (TypeSafe)Laya (open source)OpenAI Decisions API
ArchitectureAutoregressive generationParallel scoring of typed answersNon-autoregressive parallel scoringNot disclosed; fixed answer sets
Latency0.5–2+ s~150 ms median (CMU)~33 ms (self-reported)~150 ms (claimed)
Marginal cost$12.18 / 1k judgments (CMU)$0.044 / 1k judgments (CMU)$0 (self-hosted)Unpublished
Calibrated confidenceInformalYes, cascade-provenSelf-reportedUnknown (preview)
DeploymentAPIAPI, gateway-routedSelf-hosted, Apache 2.0OpenAI API (preview)

What it’s good for

The use cases are exactly the volume jobs that production systems run millions of times a day:

  • Classification and routing — which queue gets this ticket, which team owns this email, is this content policy-violating.
  • Agent next-step selection — the agent loop’s “what do I do now” branch point, where the valid actions are a finite list.
  • Tool-call gates — should this agent be allowed to call this tool on this input, yes or no.
  • Model routing and cost control — a cheap first pass that decides whether a request deserves a frontier model.
  • Triage and prioritization — scoring incoming items against a bounded severity scale.

The mental model to hold: decision = f(context, question, allowed_answers). “Which queue should receive this ticket?” is a good Decisions API question. “Write the best response to this customer” is not — that belongs in a generative model. The API is for bounded questions, and the discipline of asking whether your question is bounded is itself useful design hygiene.

The competitive read

Three players, three business models, one category. OpenAI’s bet: decisions should be a primitive inside the ecosystem you already pay for. Jev’s bet: decisions should be a specialized, calibrated, optimized service. Laya’s bet: decisions should be infrastructure nobody can charge you rent for.

The honest analysis is that OpenAI’s advantage here is not architectural — it’s distribution. A routing call inside your OpenAI stack doesn’t need a second vendor, a second invoice, or a security review of a new supplier. When a category exists only as a startup product, enterprises file it under “interesting but exotic.” When OpenAI productizes it into the dashboard those buyers already use, it graduates to default infrastructure.

But OpenAI entering also legitimizes the specialists. Jev’s moat is calibration with an evidence trail: CMU’s study found that at a 0.9 confidence threshold, Jev-accepted answers run within a point of GPT-6 accuracy, and a cascade escalating the uncertain tail kept 91.3 percent pooled quality versus 91.7 percent for GPT-6 alone at 47 percent of the fee. TypeSafe charges $0.042 per million input tokens with free output — a price line OpenAI’s margin structure has never matched. Laya is free, runs at ~33 ms self-reported, inside your own perimeter, and no cloud API will ever satisfy the compliance team that says data cannot leave.

The likely endgame is a three-layer split: OpenAI as the default layer, Jev as the optimized-cost layer, Laya as the sovereign self-hosted layer. And the winning production pattern is model-agnostic either way: a cheap, confident first pass, escalating only the uncertain tail to a reasoning model.

What to watch

OpenAI has left the decisive numbers unpublished. When broad rollout lands, three things determine adoption: whether the confidence scores are genuinely calibrated (not “contact-info style” labels), where per-call pricing settles relative to Jev’s $0.044 per thousand judgments, and whether label-set changes are truly free or require re-tuning. Meanwhile, the preview is still changing — teams wiring it into critical paths should verify the current API contract rather than trusting early writeups, this one included.

The bigger signal is architectural. Generation was the primitive of the last four years of LLM development. OpenAI shipping decisions as a first-class API primitive — alongside the 20 other DevDay launches built on agents that must constantly choose — is a quiet admission that the future workload is not more text. It is more choices, made faster, at higher volume, for less money.