One GPU, 262K Context, Apache 2.0: Agnes-3.0-Flash Makes Hybrid-Attention Multimodal Reasoning a Single-Card Affair
Agnes-AI open-sources a 33B hybrid-attention multimodal model with a 262,144-token context, image and video understanding, and 85.05 GPQA Diamond — all on one H100 at bf16.
One GPU, 262K Context, Apache 2.0: Agnes-3.0-Flash Makes Hybrid-Attention Multimodal Reasoning a Single-Card Affair
The gap between “open-weights” and “actually runnable” has been the quiet disappointment of the open-model era. Too many celebrated open releases turn out to need a node of GPUs just to load, quietly re-centralizing capability in the hands of whoever can afford the rack. Agnes-AI’s newest drop, published quietly on Hugging Face on September 11 and surfacing as the top story in aggregator feeds on the morning of September 12, is a pointed counterexample: Agnes-3.0-Flash is a 33B-parameter multimodal reasoning model with a 262,144-token context window, image and video understanding, tool calling, and an Apache 2.0 license — and the entire checkpoint fits in bf16 on a single NVIDIA H200 (141 GB) or H100 (80 GB).
What shipped
Strip away the branding and the spec sheet reads like a wishlist for local deployment:
- 33B parameters, hybrid-attention decoder, 72 layers
- 262,144-token context window — a quarter-million tokens
- Text, image, and video understanding through a 27-layer vision tower
- Adjustable reasoning effort (
high/medium/low, plus a thinking-off switch) - Native tool calling rendered straight from the chat template
- Apache 2.0 — commercial use, fine-tuning, and redistribution all permitted
- ~66 GB on disk for the bf16 checkpoint; 128 GB+ host memory recommended
That combination — frontier-adjacent reasoning, quarter-million-token context, multimodality, and a genuinely permissive license on one card — is what makes this release notable rather than merely incremental.
The architecture: three recurrent layers for every attention layer
The most interesting number in the model card isn’t a benchmark. It’s the ratio 3:1.
Agnes-3.0-Flash is a hybrid-attention decoder in which 54 of its 72 layers run a gated delta rule — a recurrent mechanism whose per-layer state size is constant regardless of sequence length — and only 18 layers run standard global attention. In other words, just one in four layers holds a KV cache that grows with context.
This is the architecture family that recurrent-and-attention hybrids (RWKV, Mamba-2 derivatives, gated delta networks) have been promising: the recurrent layers carry long-range state at fixed memory cost, while the sparse attention layers preserve the precise retrieval that pure-linear models struggle with. For a model whose headline feature is a 262K context window, the design isn’t academic — it’s the enabling decision. A full-transformer model at 33B would spend a large fraction of an 80 GB card’s memory on KV cache alone at that context length; here, three-quarters of the layers simply don’t have one.
The details are equally deliberate. The delta-rule layers use 16 key heads / 48 value heads at head dimension 128, a causal convolution (kernel 4) in front, gated RMS-norm, and fp32 recurrent state for numerical stability over long horizons. The global-attention layers use grouped-query attention (24 query heads / 4 KV heads, 6:1), head dimension 256, RMS-norm on queries and keys, and sigmoid-gated output. Feed-forward is SwiGLU with an intermediate size of 17,408 plus a parallel SwiGLU branch of width 2,048 in every layer. Positions come from a 3-axis rotary embedding (text / height / width, interleaved mrope sections 11:11:10, base 1e7) applied to the first 25% of each head dimension — the rope scheme that lets one model index tokens, image rows, and video frames coherently. Vocabulary is 248,320.
The vision tower is a 27-layer, hidden-size-1152 patch-16 design with 2×2 spatial merge, projected into the 5,120-wide decoder hidden size.
Benchmarks: honest framing, competitive numbers
The model card does something unusual for the genre: it leads with a disclaimer. The reference figures were “compiled from different sources, harnesses, and model snapshots and do not constitute a controlled head-to-head comparison” — a caveat more benchmark tables in this industry should carry, and one worth respecting when reading the columns.
With that said, the numbers are competitive for the class:
- GPQA Diamond: 85.05 — ahead of Qwen3.6-35B-A3B (84.1) and Kimi K2.5 (78.9), though behind the 90+ tier of DeepSeek V4 Flash, Qwen3.8, Gemini 3.5 Flash, and MiniMax M3
- IFBench: 74.20 — competitive instruction-following, with Qwen-family models ranging 75–83
- AA-LCR: 68.33 — long-context reasoning ahead of Qwen3.6-35B-A3B and Kimi K2.5
- SciCode: 38.08 — mid-pack, behind DeepSeek V4 Flash’s 50.3
- AA-Omniscience Accuracy: 23.00 — ahead of every comparison model except Gemini 3.5 Flash (51.4) and DeepSeek V4 Flash (40.4)
The honest read: Agnes-3.0-Flash is not claiming frontier superiority, and on raw knowledge-heavy evals the largest models retain a clear edge. The claim it can make is efficiency-adjusted — these numbers arrive from 33B parameters and one GPU, against comparison columns that include a 1T-parameter MoE (Kimi K2.5), a 284B MoE (DeepSeek V4 Flash), and a 428B MoE (MiniMax M3). On tokens-per-dollar and capabilities-per-gigabyte, the proposition is sharp.
Built for agentic work
Two features push this beyond a pure research artifact.
First, adjustable reasoning effort. The chat template exposes high (default), medium, and low levels plus enable_thinking=False, letting a deployment trade latency for depth per request — high for hard reasoning passes, low for latency-sensitive traffic. The recommended sampling settings (temperature 1.0, top_p 0.95, top_k 20, max_tokens 2000+) ship as the checkpoint’s own generation config.
Second, native tool calling. The chat template renders tool definitions directly, and the model emits structured <tool_call><function=…><parameter=…> blocks, with results fed back as tool role messages. Over an OpenAI-compatible API, tools work the same way — and Agnes ships an SGLang serving recipe (a stock nightly container plus a three-file overlay) that puts the whole thing behind a standard /v1 endpoint, --tp 2 supported for maximum context and concurrency.
Add the 262K window and video understanding, and the profile is unmistakably aimed at the agent builder who wants a self-hosted backbone: long transcripts, screen recordings, multi-step tool loops, all on infrastructure that fits in a single-GPU budget.
The catch: remote code
The release is not without friction. The checkpoint requires trust_remote_code=True — it ships its own model implementation rather than relying on merged upstream transformers support (tested on transformers 5.12.1, with torchvision required for image/video processing). For security-conscious shops, running untrusted model code is a real consideration, and it also means ecosystem tooling (quantization forks, alternative runtimes, inference engines beyond the provided SGLang recipe) will need to catch up. The Apache 2.0 weights themselves are unambiguous; the integration story is just early. History says merged upstream support tends to follow for models that get traction.
Why it matters
The open-weights movement has spent two years oscillating between two failure modes: models too small to matter and models too big to run. The interesting middle — genuinely capable, genuinely multimodal, genuinely long-context, genuinely one-GPU — has been thinner than the hype suggests. Agnes-3.0-Flash lands squarely in it.
For startups, the economics are direct: one H100 at spot rates serves a quarter-million-token multimodal reasoning model with no per-token API bill and no vendor lock-in. For researchers, a 3:1 delta-rule-to-attention hybrid at this scale with permissive weights is a usable testbed for the recurrent-attention hybridization question that a lot of architecture work is converging on. For the broader ecosystem, every release like this narrows the practical gap between the API oligopoly and what a well-funded hobbyist can self-host.
The benchmark table won’t crown it king, and it shouldn’t. But “85 on GPQA Diamond, 262K context, video understanding, tool calling, Apache 2.0, one card” is a sentence that wasn’t true last week — and that is exactly what progress at this layer looks like.