← All posts / Models

33ms, Open Weights, Better Scores: Laya Answers the Frontier Lab That Rediscovered His Idea

A solo researcher who published non-autoregressive decision models in March 2025 open-sources Laya — a 421M Apache 2.0 model that beats TypeSafe's closed Jev on accuracy, calibration, and latency.

33ms, Open Weights, Better Scores: Laya Answers the Frontier Lab That Rediscovered His Idea

One of the most hyped model categories of late 2026 is not a chatbot. It is the opposite of one: small, non-autoregressive “System 1” decision models that never generate text, never hallucinate, and return typed answers with calibrated probabilities in a single forward pass. The frontier lab that kicked off the current wave — TypeSafe AI, founded by ChatGPT co-inventor Diogo Almeida — shipped its Jev model in September as a closed API at $0.042 per million input tokens, with no papers, no weights, and no training data.

This week, the open-source community answered. Convai Innovations released Laya, a 421M-parameter multilingual System 1 decision model under Apache 2.0 — and the release comes with a personal story attached. Its author, Nandakishor M, published the same core idea on arXiv in March 2025 (arXiv:2503.23303), followed by a framework paper for schema-based decisions guided by reinforcement learning in September 2025 (arXiv:2510.01237) — roughly a year before a well-funded lab launched the concept as if it were new. Instead of staying bitter, he rebuilt the architecture from scratch, fixed its limitations, and published everything.

What Laya actually does

The pitch targets a real inefficiency in modern AI stacks: we burn generative LLMs on reflex decisions. When a support ticket, an email, or an API prompt arrives, you usually need a handful of simple answers — which department handles this, how urgent is it, does the user threaten to cancel. Calling an 8B or 70B generative model for that means 500–2,000ms of streaming, real inference cost, a fragile JSON parser on the other end, and “confidence” numbers that are just tokens that sound confident, with no mathematical calibration behind them.

Laya accepts a state (text, email, ticket, or JSON) plus typed questions, and answers all of them in one forward pass. It exposes exactly three primitives:

  • choice — picks an option and returns a probability for every option
  • score — places the input on your ordinal rubric, with a distribution
  • noul — returns a calibrated P(true) for a yes/no claim

Because the output space is strictly probabilities and numbers, there is nothing to parse and nothing to hallucinate. Broken JSON is physically impossible.

Architecture: an encoder, not a decoder

The backbone is ModernBERT-large — 395M parameters, bidirectional, fully fine-tuned — topped with a decision head trained from scratch: two transformer layers, an option-marker scorer, and an act/escalate head, for 421M parameters total. Every answer option gets its own [MASK] token; the model gathers hidden states at those marker positions and softmaxes over each question’s options. The clever consequence: the answer space is defined at request time, so entirely new schemas need no retraining.

The repo actually ships three checkpoints. The English laya (421M, 512-token context) sits at the root; laya-multilingual (322M mmBERT-base, 1024-token context, 100+ languages, roughly 2× faster) and laya-typed-decisions (421M, 1024 tokens) live in subfolders. A Router class lazy-loads and dispatches between them by input language, with an optional preload mode that the authors measured as worth up to 4.8× throughput on mixed-language traffic.

RLCD: reinforcement learning that rewards honesty

The training method, RLCD (Reinforcement Learning for Calibrated Decisions), is the most interesting part. Naive RL destroys calibration: the policy gradient pushes the winning option toward probability 1.0 and everything else toward 0.0, producing a confidently wrong machine. RLCD instead uses a strictly proper scoring rule as the reward — log score plus spherical score, with ranked probability score subtracted for ordinal questions. Expected reward under such rules is maximized only by reporting honest probabilities. Exploration adds zero-mean Gaussian noise to the logits; updates are REINFORCE with a group-mean baseline, GRPO-style. Training took 7,313 updates, one epoch, about 1.96 hours.

There is also an act/escalate head that reads the pooled [CLS] token plus distribution features and outputs P(act) and P(escalate) — the model learns when to hand off to a human. In practice, setting an automation threshold of confidence ≥ 0.85 lets Laya auto-route roughly half of incoming tickets at 92%+ precision, escalating the genuinely tricky cases.

The numbers against Jev

On a Tesla T4, a single question returns in 32.8–39.5ms; batched, Laya sustains 103–332 questions per second. Against Jev 1.13.0’s third-party published figures (TypeSafe’s API was not directly measured), routed Laya wins across the board:

  • typed-decisions accuracy: 0.766 vs 0.727
  • AG News: 0.950 vs 0.910
  • DAIR Emotion: 0.595 vs 0.480 — Jev assigned zero probability to the true label on 16% of examples
  • calibration (ECE): 0.081 vs 0.246 — 3× better
  • p50 latency: ~33ms vs 236–276ms — 6–7× faster
  • license: Apache 2.0 vs closed API; cost $0 self-hosted

The honest part

What makes this release credible is the Limits section, which is unusually blunt. Zero-shot on typed-decisions, the base checkpoints score below the majority-class baseline (0.362 against 0.461) — the headline 0.766 belongs to the checkpoint fine-tuned on that benchmark’s own training split. Laya is a fast base to specialize, not a zero-shot decision engine. Ordinal score questions are the weakest primitive. The model ships overconfident and needs a per-question-type temperature refit (which moved mean ECE from 0.466 to 0.081). And the multilingual checkpoint clears 3× random on only 23 of 51 tested languages — Khmer scores 0.000 accuracy at 0.952 confidence, a detail almost no lab would publish.

That transparency, plus Apache 2.0 weights, a PyPI package (pip install laya), a live demo, fine-tuning notebooks, and the full benchmark breakdown, is exactly what the closed alternative refuses to offer. The lesson of the week: when a frontier lab rediscovers your research and charges for it, the best reply is weights, benchmarks, and an honest README.