Yutori's Navigator n2: The 27B Model That Out-Computers Frontier Giants at One-Tenth the Price
Ex-Meta AI leaders at Yutori shipped Navigator n2, a 27B computer-use model that scores 65.2% on OSWorld 2.0 — beating GPT-5.6 Sol — for $0.50/$4 per million tokens, and tops MyPCBench by 20 points over Claude Opus 4.8.
The most interesting number in AI this week is not a parameter count. It is $1.46 — the API cost for Navigator n2, a 27-billion-parameter computer-use model from startup Yutori, to complete a full long-horizon task on OSWorld 2.0, one of the hardest agentic benchmarks in existence. On that benchmark’s partial-score metric, n2 reaches 65.2%, ahead of OpenAI’s flagship GPT-5.6 Sol at 62.6% and within striking distance of Claude Opus 5 at 70.6% — a model that costs roughly ten times more per input token.
The release, announced August 26 on Yutori’s blog, is the clearest demonstration yet of a pattern that defined the second half of 2026: computer use — agents that operate real desktops, browsers, and terminals — is becoming the new competitive frontier, and specialized mid-scale models are beating general-purpose giants on cost-performance curves that frontier labs cannot match.
What Navigator n2 actually does
Yutori’s earlier models, Navigator n1 and n1.5, specialized in browser use. n2 is a general-purpose computer-use model that operates full desktop environments across Linux, macOS, and Windows. It is trained to interleave every interface a computer offers: clicking through GUIs when visual interaction is needed, dropping to the shell for system operations, calling tools for structured data, and writing short snippets of code when that is more efficient.
That last point is the company’s core design argument, and it is worth quoting directly: “not forcing models to use computers exactly the way humans do, but enabling them to intelligently use everything a computer can do.” Renaming 100 photos from IMG_1337.jpg to paris-trip-001.jpg is one for loop in a shell — or a hundred rounds of right-click, rename, type, enter through a GUI. A model that can choose the right interface at each step finishes tasks faster, cheaper, and more reliably than one locked into emulating a human with a mouse.
The benchmark table Yutori published makes the case across five suites:
- OSWorld 2.0: 65.2% — versus Claude Opus 5 at 70.6%, GPT-5.6 Sol at 62.6%, Gemini 3.7 Flash at 47.9%, Muse Spark 1.1 at 47.3%, and GPT-5.6 Luna at 45.6%. OSWorld 2.0, released in June by the XLANG lab, shifted evaluation to complete long-horizon workflows where agents must carry information across applications.
- OSWorld-Verified: 85.3% — narrowly ahead of Claude Fable 5 at 85.0% and GPT-5.6 Sol at 83.0%.
- MyPCBench: 82.6% — a blowout. Claude Opus 4.8 manages 62.0%, GPT-5.6 Sol 55.4%. MyPCBench measures personalized computer use across a user’s own apps, accounts, and digital history.
- MacAgentBench: 83.1% — real macOS tasks spanning 25 applications. GPT-5.5 is next at 66.7%.
- WeaveBench: 70.3% — tasks requiring agents to combine GUI interaction with CLI and code execution, where Claude Opus 4.7 scores 53.2% and Gemini 3.1 Pro just 22.3%.
On OSWorld 2.0, the cost-per-task gap compounds the score gap. At $0.50 per million input tokens ($0.05 cached) and $4.00 per million output tokens, n2 sits in a different economic class from the frontier models it trades blows with. Claude Opus 5 launched July 24 at $5/$25 per million tokens. Running hundreds of agentic steps per task multiplies that difference into real money at production scale.
The recursive training loop
The most technically interesting part of the release is how the training data was made. Yutori’s insight: data generation for computer use is itself a computer-use task.
Instead of hand-authoring tasks from instruction manuals, the team used computer-use agent rollouts to explore native applications and determine what is actually feasible in each environment, grounding new tasks in real application logic. For every task, they generated a corresponding verifier — a test that inspects the end state of the environment rather than trusting the agent’s self-report. Some verifiers are programmatic: query the database, diff the file, check that the calendar invite exists. Others are LLM-based rubrics for tasks where success involves softer logic.
Then comes the adversarial pass. Candidate tasks and verifiers are stress-tested with more agent rollouts to surface false positives (wrong behavior, verifier approves), false negatives (correct behavior, verifier rejects), and reward hacks (right outcome achieved through unintended means). Each failure becomes a new test case, and tasks, verifiers, and even the mocked environments get refined. Over two months this produced more than 10,000 tasks across hundreds of applications — productivity, communication, engineering, data analysis, creative production, and scientific workflows.
The loop is recursive in the strong sense: better computer-use models help create better training data, which trains better models, which surface deeper failure modes for the next round. n2 is trained with supervised fine-tuning followed by reinforcement learning, with a continuously refreshed RL set — tasks that become reliably solvable graduate into SFT, while RL shifts toward harder tasks and newly discovered weaknesses. The training emphasizes long-horizon planning, dynamic interface selection, and self-context compaction, letting the model handle tasks requiring hundreds of steps.
Yutori also reports early success with on-policy self-distillation (OPSD), which reached peak performance four times faster in wall-clock time than group-rollout-based RL and was especially efficient at learning from tasks where the current policy had zero success — effectively letting the data frontier run ahead of the model’s capabilities.
Who is Yutori, and why the pedigree matters
Yutori was co-founded in 2024 by Devi Parikh, Dhruv Batra, and Abhishek Das — former Meta AI leaders who previously led the tech’s most-watched multimodal research. The company raised a $15 million seed in March 2025 and has been quietly building agent infrastructure since: browser-use with infrastructure included, research agents that use the web rather than just search it, always-on “Scouting” agents that monitor the web, and a “Local” product for securely logging agents into websites.
The Scout background shows in n2’s design. A company whose agents live in browsers and desktops all day cannot afford frontier-model pricing for every step, and it cannot tolerate agents that get stuck persisting with a suboptimal interface. OSWorld 2.0’s own analyses — which Yutori cites — found that today’s frontier models still struggle exactly there: either staying stubbornly on a suboptimal interface or bouncing between them without progress. n2’s training loop targets that failure mode directly.
What it means
Two implications stand out. First, the computer-use leaderboard is no longer a frontier-lab duopoly. A seed-stage startup with a 27B model — a size that runs comfortably on a single node — is publishing scores within five points of Claude Opus 5 on the hardest benchmark and beating everything else on the personalized-use suites. The capability frontier is flattening faster in agentic domains than in chat, likely because computer use rewards focused training data and verifier quality more than raw scale.
Second, the pricing resets what computer-use products can cost to run. At $1.46 per OSWorld task, an agent that operates a user’s actual machine becomes viable for consumer and SMB products, not just enterprise automation budgets. Yutori offers n2 through its API with an OpenAI Chat Completions-compatible interface, which lowers switching friction for teams already building on the standard agentic stack.
The caveats are the usual ones. Most benchmark numbers are company-reported, though OSWorld-Verified and OSWorld 2.0 are externally reproducible. n2 is not open-weights — API only, unlike the open-source waves from Zhipu, Tencent, and DeepSeek this month. And MyPCBench and WeaveBench are newer, less battle-tested suites. But the shape of the result — a small specialized model redrawn the cost-accuracy Pareto frontier of an entire capability class — is the story of 2026 in miniature, and it arrived this week from a company most people had never heard of.