The Model That Helped Build Itself: NaiveAI's First Release Is a 309B MoE With No Full Attention
The Beijing stealth startup is out of stealth: Naive-N0.5-Flash is an MIT-licensed 309B-parameter MoE with 1M-token context, zero full-attention layers, and an AI-run R&D pipeline that served 10 million sandboxes a week to build it.
Nine days after this blog covered its $1.4 billion valuation, Beijing’s most secretive AI startup has finally shipped something you can download. On September 27, 2026, NaiveAI — the company founded in February by Tsinghua associate professor Jifeng Dai — released Naive-N0.5-Flash, an open-weight 309B-parameter mixture-of-experts model built for coding and AI research. The weights carry an MIT license, the context window is a native one million tokens, and the architecture does something almost no frontier model does: it contains no full-attention layers at all.
The release turns a nine-month funding legend into a testable artifact. Until today, everything about NaiveAI was secondhand — three rounds totaling roughly $400 million, investors including Tencent, IDG Capital, and HSG, and a one-sentence website promising “100x Intelligence for the pioneers.” Now there are benchmark tables, a Hugging Face repository, an inference runtime, and a research post describing, in unusual detail, how the company believes frontier models should be built: by AI systems doing the engineering while humans set the direction.
The model: 309B total, 15.5B active, nothing but local and sparse attention
Naive-N0.5-Flash is a mixture-of-experts design with 309 billion total parameters and 15.5 billion active per token, putting it in the same weight class as Xiaomi’s MiMo-V2-Flash and DeepSeek’s V3 lineage. The architectural signature is its attention stack: 39 sliding-window attention layers interleaved with 9 lightweight DeepSeek Sparse Attention (DSA) layers in a predominantly 5:1 ratio, with grouped-query attention (GQA4) throughout. Every layer in the network is local or sparse — there is not a single global-attention layer in the model.
That choice is a bet on long-context economics. Full attention scales quadratically with sequence length; at a million tokens it becomes the dominant cost of serving a model. By keeping the entire network local or sparse, Naive-N0.5-Flash makes 1M-token contexts cheap enough to be the default rather than an enterprise upsell. The tradeoff, historically, has been recall over very long ranges — exactly what sparse attention schemes like DSA exist to recover. NaiveAI trained the model on 3.25 trillion tokens and says the research and engineering pipeline that produced it was “substantially executed by AI systems,” with humans setting objectives and evaluation standards.
The claims: cheap, fast, and built differently
The API pricing reads like a typo: $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cache reads. At those rates the model undercuts most of the open-weight field by an order of magnitude. NaiveRT, the company’s custom inference stack, is the reason it can contemplate them. It combines mega-kernel fusion, Programmatic Dependent Launch (PDL) to overlap kernel boundaries, and speculative decoding with a fused DFlash draft model, delivering 50 tokens/s per user in Standard mode and up to 2,000 tokens/s per user in Ultrafast mode. In its RL-oriented configuration, NaiveRT hits a peak single-stream decode rate of 2,122 tokens/s on 8 GPUs and cuts a full speculative decoding round from 12.3 ms under SGLang to 3.4 ms — a 72.4% reduction.
The benchmark section is dense with comparisons against GLM-5.3, Kimi-K3, Qwen-3.8-Max, DeepSeek-V4.1-Flash, Opus-5.5, and GPT-5.6-Sol across SWE-Bench Pro, DeepSWE v1.1, Terminal-Bench 2.1, ALE-CLI, and MLE-bench-30. One entry stands out as a category error turned proof point: NanoGPT SpeedRun, the competition where entrants race to train a GPT-2-class model as cheaply as possible. A model that helps humans train small models faster is the kind of capability you only build if you expect your model to be used, eventually, on itself.
The pipeline is the product
The most consequential part of the release isn’t in the model card. It’s the description of the process that built it. NaiveAI says it “broke with human-centered R&D from day one”: AI models write the code, run the experiments, monitor progress, analyze results, and iterate, while human researchers set direction, define constraints, and make critical decisions. The infrastructure supporting this serves close to ten million sandboxes per week, with 100,000 active concurrently at peak, all managed through a unified control plane handling compute, environments, tools, permissions, and security.
The NaiveRT case study quantifies what that looks like in practice. The runtime was built in six days through 151 documented optimization trials — 43 on whole-network engineering (28 adopted), 45 on real-checkpoint execution (15 adopted), and 63 on W8A8 kernel refinement (20 adopted). Sixty-three of the 151 trials produced adopted changes; 71 failed validation or were rolled back; 17 were exploratory. The post is candid about the failures in a way most lab blogs aren’t: a MoE fusion path went through seven implementation rounds, each numerically correct, each slower end-to-end than the PDL-chained baseline, until human researchers killed the direction. One DSA layer went from 29 kernel executions per decoding step to a single cooperative mega-kernel running across 148 CTAs. A single QKV weight-prefetch optimization — slower in isolated microbenchmark, 21–23 μs faster per step in the full model — was kept.
That last detail is the quiet argument of the whole release. The AI system judging candidate optimizations by end-to-end impact rather than microbenchmark wins avoided the classic trap of kernel engineering, and did so across hundreds of trials at a pace humans don’t sustain. Six days for a custom inference runtime that beats SGLang’s decode latency by 3.6x on their target workload is the kind of number that either validates the AI-centered R&D thesis or will be picked apart in replication.
A world model built by the model
The second case study goes further afield. A NaiveAI researcher handed Naive-N0.5-Flash an end-to-end research task outside the company’s expertise: build a world model. The human defined the objective, compute budget, and evaluation protocol; the model designed approaches, trained models, ran evaluations, and decided what to try next. After 400 cumulative hours and 15 major experimental rounds, the resulting model — AutoWM — scored 77.43 on the public WorldArena-1 Track 1 protocol, above the then-highest published leaderboard score of 73.64. The model had begun by reproducing the official FlowWAM recipe, found conventional tweaks unhelpful, and moved to rewriting the problem itself: regenerating captions and expanding the training set from 2,500 to 22,500 clips.
Because Naive-N0.5-Flash is also trained for AI R&D work — participating directly in the research loop — NaiveAI frames the release as “opening a path toward recursive self-improvement.” That phrase does a lot of load-bearing work, and it’s the one to watch. Recursive self-improvement is the theoretical engine of fast takeoff scenarios: systems that improve the systems that improve them. A 309B open-weight model with strong agentic coding skills, a million-token context, sandbox infrastructure already operating at ten million jobs a week, and an explicit design goal of participating in its own successor’s development is either the most efficient research lab ever open-sourced or the most interesting unaligned agent ever shipped under MIT license — and the honest answer is that nobody yet knows which.
What it means
Three things make this release more than another open-weight drop. First, the economics: 15.5B active parameters and $0.10/M input pricing put frontier-adjacent coding capability within reach of any team with a single node and a Docker Compose file — NaiveRT ships with one, and the full source plus W8A8 checkpoint lands by October 12. Second, the architecture: a no-full-attention design at 309B scale with a native 1M context is a data point that the sparse-attention path DeepSeek blazed is becoming the default for long-context serving. Third, the process: if the AI-centered R&D claims hold up under community scrutiny — and with open weights and an open runtime, they can be checked — the cost of building competitive models just dropped for everyone, not just NaiveAI.
The company that was a one-sentence website and a $1.4 billion valuation nine days ago is now a shipping competitor to GLM, Kimi, Qwen, and DeepSeek, with a methodology that challenges how all of them staff their labs. Dai’s bet from the beginning was that mid-training and post-training over open bases held more headroom than pretraining arms races. Naive-N0.5-Flash is the first installment of that bet made public — built, the company says, substantially by the systems it is meant to help create.