88.6% on BrowseComp, Weights Promised: AllSpark's Iris Agents Take the Open-Source Search Crown
AllSpark's Iris-mini and Iris-pro open-weight search agents set the pace for open models on BrowseComp, DeepSearchQA and HLE, with a fully documented SFT-RL climbing recipe.
Open-source search agents just got a new reference point. A team calling itself AllSpark has published “Iris: Climbing to the Search Frontier” (arXiv:2609.04304, submitted September 3, 2026), a 12-page report describing Iris-mini and Iris-pro — two open-weight search agents that now post the strongest overall results among open-source systems in their parameter ranges on four of the hardest agentic search benchmarks in circulation. Within days, the paper had become the top-ranked story on AI news trackers, and the reason is simple: it is not just a leaderboard grab, it is a full recipe.
What was released
Iris-mini is a 35B-A3B mixture-of-experts model initialized from Qwen3.6-35B-A3B; Iris-pro is a 397B-A17B model initialized from Qwen3.5-397B-A17B. Both keep the 256K-token context window of their base models. The “A” in those size labels matters: only 3B and 17B parameters are respectively active per token, which is what makes a 397B-parameter search agent remotely serviceable outside a hyperscaler’s training cluster.
The headline numbers, with context management enabled and evaluated as a single ReAct agent with no sub-agents and no test-time verification:
- Iris-pro: 88.6 on BrowseComp, 85.1 on BrowseComp-ZH, 92.9 on DeepSearchQA (F1), and 56.4 on the text-only subset of Humanity’s Last Exam.
- Iris-mini: 82.2 on BrowseComp, 84.8 on BrowseComp-ZH, 86.9 on DeepSearchQA, and 52.3 on HLE.
For context on how sharp the open-source field has become: in the 30–35B class, Iris-mini beats the previous best, XYZ-Aquila-mini, by 3.4 points on BrowseComp (82.2 vs 78.8). In the ~400B class, Iris-pro outperforms XYZ-Aquila-pro by 3.8 points on BrowseComp and 3.1 points on HLE, ties it on BrowseComp-ZH at 85.1, and edges DeepSearchQA 92.9 vs 92.5. The team also states that Iris-mini approaches 1T-scale models such as Kimi-K2.6 and DeepSeek-V4-Pro on BrowseComp — an 83.2 and an 83.4 in the paper’s comparison table, against Iris-mini’s 82.2 with roughly a thirtieth of the active parameters.
The gap to the frontier has not closed. GPT-5.6 Sol sits at 90.4 on BrowseComp, Kimi-K3 at 91.2, and Claude Fable 5 posts 64.5 on full-set HLE and 94.2 on DeepSearchQA. Iris is the king of the open weights, not of search overall — a distinction the authors themselves are careful to make.
The interesting part: how the training data is built
Leaderboard numbers age quickly; the data pipeline is what makes this paper worth reading. The core problem in training search agents is that naturally occurring web questions are too easy — a model can often answer them from parametric memory without searching at all. Several recent systems have attacked this by traversing hyperlink or knowledge graphs and masking entities along the path. Iris pushes that idea into a fully automated, LLM-driven pipeline with three stages.
Web-graph construction. The corpus is modeled as a directed graph: pages are nodes, hyperlinks are edges. A seed page is chosen in “answer-anchored” mode — fix a target answer entity, retrieve pages about it — and expanded along its out-links into a local subgraph with full page text retained up to a budget.
Task synthesis. The subgraph is distilled into a compact entity graph of salient entities and typed semantic relations, keeping only entities that lie on a multi-hop path toward the seed theme. A question is then authored over that graph with a hard constraint: the reasoning path must traverse at least N coupled relations, and the answer must not appear in the question. Then comes the cleverest step, anchor abstraction: every non-answer entity mentioned in the question is rewritten into a descriptive reference that uniquely identifies the entity but never uses its name or aliases. The result is a question that cannot be cracked by string-matching a surface form into a search box — the agent has to disambiguate by reasoning.
Dual-criteria verification. A question is admitted only if a reference model fails it closed-book (proving it genuinely requires search) yet solves it when the supporting evidence graph is supplied (proving it is solvable and has a unique answer). Both checks are decided by a semantic-matching judge. Only the intersection survives.
SFT-RL climbing, and honesty about the harness
The training recipe alternates supervised fine-tuning with reinforcement learning against live web search. Teacher trajectories over the accepted questions are filtered twice — at the trajectory level for correctness, degeneracy, and search depth, then at the turn level by a judge whose rubric is induced from the data rather than hand-written. RL runs on the open-source Relax framework, with the reward judge and observation summarizer served inside the training cluster, and with over-long rollouts interrupted at the request level and resumed from their committed prefix.
The alternation is what the team calls “SFT-RL climbing”: the hardest solved and most efficient rollouts of each RL round are fed back into the next supervised pass, so behaviors discovered during exploration get consolidated before the next round begins. It is a bootstrapping loop with a ratchet.
Just as notable is the evaluation hygiene. The paper’s central methodological claim is that inference-time context management — summarization, history compression, the “discard-all” context reset popularized by DeepSeek-V3.2 — is worth more on these benchmarks than most reported differences between systems. So every benchmark is reported both with and without context management, holding the tool set, context limit, and judge fixed. The team also blocked access to Hugging Face dataset and Space pages hosting the benchmark answers at three enforcement points, to prevent benchmark leakage from contaminating RL rollouts. In a field where “solved” often means “found the answer in the training data,” that paragraph alone deserves credit.
Why it matters
First, open-weight search capability is compounding fast. The distance between the best open model and the previous best open model in the ~400B class is measured in single digits, and a 35B-A3B model now hangs within a few points of 1T-scale systems on BrowseComp. Specialization — Iris is trained purely for web search and long-horizon information seeking — is buying performance that raw scale used to monopolize.
Second, the release matters more than the numbers. The team plans to publish the model weights together with the complete recipe for data construction, training, and evaluation. Reproducible SFT-RL pipelines for agentic search have historically lived behind closed doors; a documented, leakage-guarded, dual-verified data pipeline is the part other teams can actually build on.
Third, the context-management finding is a quiet warning for benchmark consumers. If harness choices swing scores by more than the margins between competing systems, then agent leaderboard rankings partly measure inference scaffolding, not intelligence. Iris’s with-and-without-CM reporting sets a standard other evaluations should copy.
The open questions are the usual ones: how the models behave on live, adversarial, non-benchmark tasks; what the contamination risk is for benchmarks whose answers are now themselves training data for the next generation of agents; and whether AllSpark — which describes itself only through the paper, with a motto of “Do the right things, and do things right” — follows through on the weight release. But as of this week, the frontier of open-source search has a new marker, and it comes with instructions.