64% on Terminal-Bench From a 122B MoE: How T1 Turned Agent RL Into an Infrastructure Problem
Tencent Hy's T1 post-trains Qwen3.5-122B-A10B with pure verifier-driven RL to reach 64.0% on Terminal-Bench 2.1 — and its real contribution is the unglamorous plumbing: TITO token-faithful training, R3 routing replay, and a dense assertion-count reward.
The most interesting terminal agent of the week wasn’t announced on a San Francisco stage. It arrived as an arXiv preprint, submitted September 10, 2026, from a team spanning Tencent’s Hy Foundation Model Frontier group, the National University of Singapore, the University of Georgia, Indiana University, and the University of Maryland. The model is called T1 — a 122B-parameter Mixture-of-Experts agent, built on the Qwen3.5-122B-A10B base and post-trained with PPO to operate a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task’s own held-out verifier.
The headline numbers are strong. On Terminal-Bench 2.1, T1 resolves 64.0% of tasks, lifted from a 43.8% base model — a 20.2-point climb from post-training alone. On Long-Horizon Terminal Bench, it reaches 27.9%, which the authors report surpasses GPT-5.4 (27.2%) and GLM-5.1 (26.7%) and matches Gemini 3.1 Pro. On the independently constructed Terminal-Bench Hard, T1 scores 38.0%, exceeding DeepSeek V4 Pro’s 36.0%. Under the same evaluation harness, T1’s 64.0% sits above GPT-5.4 (54.8%), DeepSeek-V4-Flash (56.9%), and Claude Opus 4.6 (63.8%), and just below Claude Opus 4.7 (66.1%) — best in its size band.
But the scores are not why the paper matters. T1 matters because of what it reveals about how agent capability is actually made in 2026: not with a novel architecture, but with an obsessive stack of infrastructure fixes that most labs don’t publish.
Three-quarters of the gain came from RL, not imitation
The pipeline decomposition is instructive. The base model starts at 43.8% on Terminal-Bench 2.1. An SFT checkpoint — initialized from demonstrations — reaches 49.4%. That is a 5.6-point contribution. Reinforcement learning on terminal tasks then adds another 14.6 points to arrive at 64.0%. By the team’s own accounting, RL earns roughly 72.3% of the total distance from base to final. Imitation got the model a fifth of the way; learning from executed outcomes got it the rest.
The reward design explains why. The team’s first campaign used a binary outcome — a task either passed its verifier or it didn’t — and it never surpassed the supervised baseline. The problem is information density: a rollout batch costs hundreds of sandbox-hours, and a binary outcome returns exactly one bit per trajectory. Tasks too hard to solve outright all score zero, even when half the requirements were satisfied.
So T1 spends the verifier’s full resolution instead. Each task’s verifier reports every assertion individually, and the reward is the absolute number of passing assertions, normalized on a scale fixed once for the entire run. The design choices are deliberate: an absolute count rather than a ratio (so half of a long task doesn’t look like half of a short one), a fixed global normalizer rather than a per-batch maximum (so the critic never chases a drifting target), and a fallback to the binary outcome when per-assertion reports are missing (so a solved task never scores zero).
TITO and R³: closing the train–inference gap on sparse models
The technically novel core of the paper addresses a problem that haunts on-policy RL for Mixture-of-Experts models. Between the trajectory a sandbox produces and the gradient the trainer applies lie dozens of turns, two independent execution stacks, and a discrete router — each able to silently attribute the update to a policy that never generated the data.
Two mechanisms attack this. TITO (token-in, token-out) makes the agent exchange token identifiers rather than text: the sampler returns token ids with their log-probabilities, encoding runs once per turn instead of once per history replay, and templates are pinned so rendering stays append-only. Where the harness still forces a mismatch, an assembler repairs turn boundaries under four escalating cases — strict, normalized, retokenized, split — and an auditor re-locates every completion inside the next prompt, reporting zero token drift inside the loss region. Critically, the harness’s re-tokenized copy never enters the training stream.
R³ (rollout routing replay) handles the router. Expert selection in an MoE is a discrete top-k, so a numeric difference far below any logit tolerance can swap a selected expert for its runner-up — substituting one sub-network for another between rollout and training. R³ records the expert sets inference selected and constrains the training forward pass to reuse them, while gating still reads live parameter scores so the router remains trainable. A missing capture aborts the rollout rather than silently dropping the sample.
Together, TITO and R³ cut the mask-weighted mean training-to-inference log-probability gap from 0.021 to 0.013 — roughly a third — with exactly aligned tokens in the loss region.
What they deliberately left off
The “withheld” section reads like a contrarian checklist. Both KL terms are disabled, because a frozen reference routes with its own expert selection and would charge the policy for a bookkeeping difference. The MoE load-balancing coefficient is zero, because balancing pressure asks the router to redistribute exactly the choices replay asks it to reproduce. No learning-rate schedule, no entropy bonus, no length shaping — the additive length penalty they tried produced unbounded trajectory growth, while the plain dense reward doubled turn counts and then flattened on its own.
The training data is equally audited. T1-15k is a 15K-task subset surviving recursive task synthesis (15 rewrite rounds, five selection stages) and an LLM audit weighted 45% verifier quality, 25% solution quality, 20% instruction quality, and 10% task value. Tasks are hard-rejected for hidden requirements, test leakage, solution shortcuts, or verifiers too weak to validate the goal — because the first defense against reward hacking is refusing gameable tasks before they enter the pool.
The caveats worth stating plainly
The paper is candid about limits. TITO is exact only up to re-tokenized history; eliminating the residual would require harnesses to treat token ids as the source of truth. The binary-versus-dense reward evidence is campaign-level, not a matched ablation. The verifier runs inside the agent-controlled sandbox with no runtime tamper detection. And oversampling — launching more trials than the batch needs and admitting the first to complete — systematically drops the hardest trials from training, precisely where evaluation failures concentrate. Weight release, cost, and latency figures are not reported.
One more number deserves attention: average turns per task ran 94.4 for T1 versus 31.5 for the SFT checkpoint. Longer interaction is both a capability and a cost — on representative unsolved tasks, T1 spends hundreds of turns where stronger baselines spend tens, and times out. The same harness-sensitivity caveat applies to the comparisons too: Claude Opus 4.6 scores 70.1% under Claude Code but 63.8% under the Terminus-2 harness T1 used, so cross-harness leaderboard rows are context, not verdict.
The takeaway for practitioners is the one AI Weekly’s editor drew: terminal-agent RL has become a shopping list of infrastructure fixes rather than modeling tricks — and a 122B recipe that methodically closes the on-policy log-prob gap is exactly the kind of detail that decides which teams can attempt this class of work at all.