One Endpoint, Three Answers: OpenAI's Decisions API Enters Public Beta on GPT-6 Luna
OpenAI's Decisions API — a dedicated POST /v1/decisions endpoint that returns typed probabilities, choices, and rubric scores about 10x faster than the Responses API — is now in public beta on GPT-6 Luna at $0.10 per million input tokens with no output charges.
Seven days after teasing it at DevDay 2026 as a limited preview, OpenAI has flipped the Decisions API into public beta. The October 6 entry in the official API changelog is terse — “Released the Decisions API in beta with gpt-6-luna” — but the move completes OpenAI’s entry into the fastest-forming new category in AI infrastructure: models that don’t write prose at all, and instead exist to answer one bounded question, fast.
The pitch, in OpenAI’s own docs: the Decisions API “evaluates text, images, or both and returns typed answers about 10x faster than the Responses API.” Get the probability that a condition is true, a selection from a fixed set of options, or a score against a rubric — and use those answers to classify content, route requests, and prioritize work. No completions to parse, no JSON to validate, no prompt engineering to keep a chatty model on task.
How it works
The API is a dedicated route — POST /v1/decisions — and a request has three parts. The model field is currently a single choice: gpt-6-luna is the only model served. The input is shared evidence for all questions: a plain text string, or user messages containing text and images. The questions array carries what to evaluate: each question’s type, instructions, and any allowed choices or score levels.
Three question types cover the intended surface:
predicate— check a condition (“does this product photo show visible damage?”). Returnsprobability, an estimate from 0 to 1.choice— select one option from a set you define (a department, a content category). Returns the selectedchoiceplus a probability distribution over options and aconfidencefield.score— rate an input against ordered levels such as issue severity. Returns the probability-weighted average of level indices — a score that can legitimately land between levels.
Independent questions can share one request — check a product photo for damage and classify its category in the same call, each with a different type. Dependent decisions require separate requests, a constraint the docs illustrate with a practical example: check for damage first, then use the result to decide whether to request a repair category. Answers come back in an answers array keyed by each question’s unique name, and a refusal answer type handles inputs the model won’t evaluate. OpenAI’s guidance is refreshingly grounded: write questions around observable criteria, give choices distinct meanings, and set thresholds using labeled examples from your own application, tuned to the real cost of false positives versus false negatives.
The docs position the boundary against OpenAI’s other structured-output machinery: use Decisions when you need one of these three answer types; use Structured Outputs with the Responses API when you need to generate an object following your own JSON schema, and function calling when a model should request a tool with arguments. There’s also a voice path — client delegation with the Live API lets developers choose actions from spoken requests and report results back — which hints at where OpenAI sees this fitting inside realtime agent loops.
Pricing and availability
The commercial model is the most aggressive part of the launch. Input costs $0.10 per million tokens, and that’s the whole bill: no cache-read charges, no cache-write charges, and — remarkably for an LLM endpoint — no output-token charges at all. When your output is a probability and a handful of enum values rather than streamed prose, charging for output tokens stops making sense, and the pricing reflects that. Regional processing premiums and long-context input multipliers still apply.
Enterprise boxes are checked from day one: Zero Data Retention and HIPAA eligibility for qualified customers, with data residency and regional processing in the United States and Europe (EEA plus Switzerland). SDK support landed across five languages — Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, and Java 4.78.0 — and there’s a Playground for experimenting with questions before writing code. OpenAI says general availability is expected “in the coming weeks.” OpenRouter has already listed the endpoint as gpt-6-luna-decisions at the same $0.10/M input, $0 output.
A category that formed in three weeks
Context matters here. TypeSafe AI — founded by ChatGPT co-inventor Diogo Almeida — shipped Jev, the first “System One” decision model, on September 15, claiming up to 200x faster inference and 400x lower cost than comparable LLMs on bounded decision tasks. OpenAI previewed the Decisions API at DevDay on September 29, a timing many read as a response. Cloudflare followed on October 1, open-sourcing Clef and Clef-flash under Apache 2.0 with median latency as low as 38.8 ms, and Perplexity launched its own Decisions API on the open-weights pplx-decider-v1-27b the same week, claiming 85.71% on an eleven-benchmark panel against Jev. AWS and others are circling.
Within that field, OpenAI’s differentiation is not raw speed — Clef-flash’s sub-40 ms medians are hard to beat from a hyperscaler’s standard regions. It’s three other things: GPT-6 Luna’s frontier-grade judgment on ambiguous multimodal inputs, native vision support in the same call, and the enterprise posture (ZDR, HIPAA, data residency) that decision-heavy pipelines in healthcare and finance actually require before they can adopt the category at all.
The Hacker News thread on the beta drew roughly 250 points, with commenters noting both the 10x speed claim over the Responses API and the view that this was “a rush job to respond to the competition.” Both readings can be true. Simon Willison shipped an llm-openai-decisions plugin for his LLM CLI tool within hours of the beta opening — a sign of genuine developer pull.
Why it matters
The decision-model wave is really an economics story wearing a latency costume. Agent frameworks burn most of their tokens not on thoughtful generation but on millions of tiny gatekeeping calls — should this ticket escalate, is this document relevant, does this step need review — each of which used to spin up a full LLM completion. Moving those calls to a typed, probability-first endpoint at a tenth of a cent per million input tokens changes the unit economics of agentic software by orders of magnitude, and it replaces brittle output-parsing with answers that are correct by construction.
For OpenAI, the strategic significance is defensive and offensive at once: it counters the upstarts on their home turf while converting the gating layer of every agent stack — the highest-volume, most commoditized call in the pipeline — into an OpenAI-billed line item. For developers, the public beta means the category now has a frontier-lab option with published pricing, five SDKs, and an availability date for GA. The three-week sprint from Jev to a public OpenAI beta is among the fastest category-to-incumbent cycles AI infrastructure has seen — and with GA “in the coming weeks,” it isn’t finished yet.
Sources
- [1] https://developers.openai.com/api/docs/guides/decisions
- [2] https://developers.openai.com/api/docs/changelog
- [3] https://openai.com/index/devday-2026-recap/
- [4] https://aiweekly.co/alerts/openai-opens-decisions-api-public-beta-on-gpt-6-luna-prices-input-at-010-per
- [5] https://news.ycombinator.com/item?id=49984025
- [6] https://openrouter.ai/openai/gpt-6-luna-decisions
- [7] https://blog.cloudflare.com/clef-decision-models/