← All posts / Models

Shanghai AI Lab Open-Sources Atria Dawn Preview: A 744B Agentic Model Built to Finish What It Starts

MIT-licensed weights, a 256K context window, and top scores on BrowseComp, DeepSearchQA, BFCL v4 and CyberGym — plus a 769-task study of how researchers actually work alongside it.

Shanghai AI Lab Open-Sources Atria Dawn Preview: A 744B Agentic Model Built to Finish What It Starts

For four days, one of the most capable open-weight agentic models ever released simply sat on Hugging Face with no announcement at all. The weights for Atria Dawn Preview were uploaded on September 11 by the Shanghai Artificial Intelligence Laboratory; only this week, with the publication of a 143-author paper titled Atria Dawn: The Dawn of Agentic Superintelligence (arXiv:2609.15818), did the model start trending worldwide. It is a preview release of a new-generation agentic model, built on the 744-billion-parameter Mixture-of-Experts GLM-5.2 foundation, with a 256K context window, MIT-licensed weights, and an unusually explicit design philosophy: don’t treat a task as complete until the output is executable, verifiable, and reproducible.

What Atria Dawn Preview actually is

Most frontier “agentic” models are chat models with tool-calling bolted on. Atria Dawn Preview is described by its creators as a foundation agentic language model — the agency is the product, not a feature. The model drives open-ended problems toward results that can be checked by running them, combining task objectives with environmental feedback to support the full loop: problem analysis, solution design, tool use, code implementation, experiment execution, result analysis, and failure recovery.

Shanghai AI Lab organizes the model’s capabilities across four dimensions:

  • Discovery — retrieving and organizing evidence, conducting deep research, and turning research questions into executable experimental plans.
  • Creation — building software, interactive applications, games, data visualizations, and machine-learning systems.
  • Delivery — transforming documents, data, and design requirements into reports, presentations, and other structured deliverables.
  • Cybersecurity — analyzing security issues, validating vulnerabilities, applying fixes, and re-validating them in authorized environments.

That fourth dimension is notable. Very few model cards treat offensive-and-defensive security work as a first-class capability alongside research and office productivity — and the benchmark numbers suggest it isn’t marketing.

The benchmarks: frontier-competitive, best-in-class on five

Across 16 benchmarks spanning search, coding, tool use, productivity, and security, Atria Dawn Preview is competitive with frontier agents — and posts the highest reported score on five of them, including several where it beats GPT 5.6 sol, Claude Opus 5, DeepSeek V4 Pro, Kimi K3, Qwen 3.8 Max, and GLM 5.3:

BenchmarkAtria DawnBest competitor
DeepSearchQA96.0Kimi K3 — 95.9
BrowseComp92.5GPT 5.6 sol — 92.2
BFCL v4 (tool use)77.0GLM 5.3 — 74.1
AutomationBench53.8Qwen 3.8 Max — 49.7
CyberGym86.5GLM 5.3 — 84.5

The deep-research numbers are the headline: 96.0 on DeepSearchQA and 92.5 on BrowseComp put it ahead of every closed frontier agent listed in the evaluation table. AutomationBench — which measures end-to-end task automation rather than single-turn tool calls — shows a similar story, with a margin of more than four points over its nearest rival and roughly eight over GPT 5.6 sol and Claude Opus 5.

It is not uniformly dominant. On SWE-bench Pro it scores 59.6 against Claude Opus 5’s 74.7; on Terminal-Bench 2.1 it manages 78.3 versus Claude’s 90.2; and on GDPval and JobBench the closed frontier still leads. In other words: state-of-the-art at finding things out and automating workflows, merely competitive at hard software engineering. The model also accepts text input only — the endpoint rejects images outright.

The Verifiable Experience Pipeline

The training method is arguably the paper’s core contribution. Atria Dawn Preview is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Rather than optimizing against rubric-graded or human-preferred responses, the model learns from trajectories whose success or failure is determined by what actually happens when the code runs and the tools fire — the same property that makes its deployment-time behavior “verifiable by construction.”

This is the same broad direction the field has been moving — RL against executable rewards rather than human preference — but applied at unusual scale and breadth, across research, engineering, and office-work environments simultaneously.

The 769-task human-AI study hiding inside the paper

The most unusual section of the paper has nothing to do with benchmarks. The team analyzed the actual development process of the model itself as a case study of human–AI collaboration: 769 task records from 56 participants, paired with agent logs. When those participants were asked to evaluate completed tasks under comparable conditions, they rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently proposed methods and implemented revisions, while humans retained most final decisions and guided exploration through judgment and feedback.

The authors’ framing is worth quoting in spirit: this marks a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. They argue that progress toward more autonomous AI research must advance both the capacity for discovery and the capacity for meaningful human oversight — preserving accountable human authority over the risks and direction of continued development. A paper about building a frontier agent that devotes its conclusion to arguing for human authority is a rare thing.

Availability and deployment

Atria Dawn Preview ships in two variants — full-precision (bf16) and FP8-quantized — both with 256K context, on Hugging Face (internlm/Atria-Dawn-Preview) and ModelScope. Local deployment is supported on SGLang v0.5.13.post1+ and vLLM v0.23.0+. Hosted access is available through an international API console at api.atria-asi.ai and a China endpoint on the intern-ai.org.cn discovery platform, and the model can be wired into OpenAI Codex as a custom provider via the Responses API.

Why it matters

Three things make this release more than another big open-weight drop. First, the MIT license on a 744B-parameter MoE with genuine frontier-level agentic scores removes almost every friction for commercial adoption — no usage caps, no regional restrictions baked into the license. Second, the provenance: Shanghai AI Lab building on Z.AI’s GLM-5.2 backbone is a portrait of the Chinese open-weights ecosystem compounding on itself, the same week Z.AI closed its own multi-billion-dollar raise. Third, the evidence-first evaluation story: a model trained to terminate only at verifiable outcomes, evaluated heavily on exactly the benchmarks (deep research, automation, security) where verifiability matters most.

The quiet upload on September 11 followed by a global trend four days later also says something about how releases work now — the weights spoke first, the paper second, and the marketing not at all. For a preview model, Atria Dawn arrives with unusually complete documentation of both what it can do and how its builders think it should be governed. The dawn metaphor may be grandiose, but the artifact underneath it is real, open, and — on at least five leaderboards — simply the best there is.