Apodex 1.1 Pitches 'Environment Scaling' for Agents and Ships a 35B Mini You Can Run Locally
A 70-author paper claims two new scaling axes — executable environments and agentic coordination — push a 35B open-weights model into frontier territory on finance and science benchmarks.
While the AI industry’s attention has been locked on raw model size and chat benchmarks, a quieter argument has been building: that the next real gains for agents will come from scaling the environments they work in, not just the parameters they think with. On August 24, a team of roughly 70 authors published “Apodex 1.1: Scaling Agentic Intelligence for Complex Work,” and the release quickly became the top research trend on Hugging Face Papers the following day. Alongside the paper, the team shipped something unusual for a research artifact: a 35-billion-parameter open-weights model, an execution harness, and a live online workbench — all available now.
The claim at the center of the release is deceptively simple. Deep-research products are good at reading web pages and writing structured reports, but real professional work — cleaning a messy clinical dataset, working through a contract file, hedging FX exposure — only begins after the report would have been delivered. Apodex 1.1 is built to work inside files, code, and tools from raw input to a verifiable deliverable, rather than summarizing from the outside.
Two new scaling axes
The paper’s core thesis is that agentic capability scales along two underexplored dimensions:
Environment Scaling. Rather than simply adding more coding data, the team systematically built training tasks around real scenarios — complex file processing, code execution, tool use, and multi-step task progression — spanning search, file, and tool environments at a scale of tens of millions of tasks. The idea is to expand the set of executable environments a model can learn to act in, on the premise that reasoning lands in the real world through tools: when a judgment needs data, the model reads files and runs code; when a premise needs verification, it searches and checks.
Agentic Coordination Scaling. The second axis scales a task’s ability to be decomposed, coordinated, integrated, and reorganized across multiple agents, branches, and points in time. In the product’s “Deep Discover” mode, the model itself — not a pre-written orchestration script — decides whether a task can be split, how to split it, and how many subagents to spawn. Intermediate results flow continuously back into a shared task state, so the main task can absorb findings and reorder priorities without waiting for every branch to finish.
Both dimensions run on a shared runtime the team calls AgentOS, which maintains tool calls, file state, and task progress, manages subagent lifecycles, and provides unified verification. The two scaling paths are explicitly not separate features bolted together — they share the same underlying runtime, which is why file-environment gains and coordination gains could be pushed forward in the same training and evaluation system.
The numbers
On evaluation, the team reports that Apodex 1.1 with its Agent Team setup achieves frontier-level scores across professional work, finance, and scientific research: 38.5 on APEX-Agents, 78.8 on GDPVal, 54.3 on FrontierFinance, 63.3 on FrontierScience-Research, 35.3 on BioMysteryBench, and 56.1 on Humanity’s Last Exam. The Agent Team configuration consistently improves over a plain ReAct setup across all six benchmarks, and posts the highest scores among compared systems on FrontierFinance and FrontierScience-Research. To guard against benchmark contamination, the team blocked access to benchmark-hosting websites during evaluation.
The more striking result is the Mini. The 35B Apodex-1.1-mini reaches the performance band of selected frontier systems on professional work, finance, and science — it leads FrontierFinance at 50.2 and lands 27.7 on APEX-Agents, close behind the full system. Because several proprietary competitors do not publish parameter counts, the team deliberately avoids unsupported size comparisons; the headline is that a disclosed 35B-scale model competes in the same complex-work tables as frontier systems.
What’s under the hood
Several technical choices stand out from the technical report and model card:
- Base model. Apodex-1.1-mini is built on Qwen3.5-35B-A3B, a Mixture-of-Experts architecture with roughly 3B active parameters, released under Apache 2.0. The NVFP4 quantized variant ships with an 8-bit mixed-precision config and a 262,144-token context window.
- PIVOT-RL. Long-horizon tasks provide only coarse terminal rewards — a successful trajectory can contain weakly grounded decisions, and a failed one can contain useful early work. PIVOT-RL runs hindsight-guided analysis over hundreds of thousands of trajectories to locate the “pivots” where a run goes off course, then constructs localized continuation tasks at those points. Corrective hints are used only during training and never appear at inference.
- Statement Review. Key claims are checked by an independent review step against sources, data, and computations before delivery. When a citation does not match or evidence is insufficient, the result is sent back for correction rather than passed through.
- AI4AI. In a telling experiment, the team used Apodex 1.1 as a fully automated teacher — writing questions, filtering trajectories, and evaluating progress with no human involvement — to teach Qwen3.5-0.8B to use search tools and authoritative databases. After 10 automated rounds, the small model’s score rose from 51.0% to 56.0% across 200 clinical-trial, protein-structure, and protein-annotation questions.
The workbench and the harness
The primary release is the full Apodex 1.1 model behind an online workbench at apodex.ai, where users can bring papers, datasets, spreadsheets, and code into one long-running task with mid-task intervention. Demonstrated use cases read less like chatbot demos and more like junior-analyst work product: a bankruptcy clawback analysis that reconstructed six preferential transfers, computed $550K net exposure, and proposed a $175K–$350K settlement range with a client-ready memo; an FX hedging task that not only priced option structures and a $7.6M net collar but caught a flawed accounting assumption and re-derived the ASC 815 treatment; and a molecular-dynamics setup that converted the 7M6J protein structure using the Martini 3 coarse-grained method for downstream simulation.
For developers, the open-source FrontierAgent harness (GitHub: ApodexAI/FrontierAgent) provides a terminal TUI with two modes — ReAct for sequential single-agent work, and Agent Team for parallel subagent dispatch. It runs out of the box on macOS and Linux without Docker, falling back to a native runtime when containers are unavailable. The team also previewed FrontierSearchBench and FrontierResearchBench, with a dedicated benchmark-release blog promised.
Why it matters
The release is a data point for a shift that has been gathering steam all year: the frontier of agent capability is moving from “bigger models” toward “richer environments and better coordination.” If a 35B MoE model with an open harness can sit in the same benchmark tables as frontier systems on finance and science work, the moat around proprietary chat models gets narrower exactly where enterprises actually pay — long-horizon, verifiable, file-heavy professional tasks.
There are caveats, of course. The benchmarks are niche and partly novel (several were introduced by the team itself), and self-reported evaluation of one’s own model always deserves scrutiny. But the full recipe — weights, harness, paper, and live product, with contamination controls stated explicitly — is more open than most frontier labs offer, and the “environment scaling” framing is likely to be copied.
Pretraining for the next generation is already underway, and the team says much of this release’s engineering and data work will carry directly into future versions. Expect the two scaling axes to become standard vocabulary.