← All posts / Models

DeepSeek V4-Flash-Vision-Exp: The Experimental Multimodal Model Chasing Opus 4.8

DeepSeek's experimental vision-equipped V4-Flash lands within points of Anthropic's Opus 4.8 on multimodal agent benchmarks — at a fraction of the price.

DeepSeek V4-Flash-Vision-Exp: The Experimental Multimodal Model Chasing Opus 4.8

DeepSeek has quietly shipped one of the more interesting model releases of the summer. On August 21, 2026, the Hangzhou-based lab took the wraps off DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal version of its flagship text-only V4-Flash model. The new model can read images and screenshots — and then act on what it sees, a combination that pushes it directly into competition with Anthropic’s Opus 4.8 on multimodal agent benchmarks.

The release is live now on the DeepSeek API platform under the model ID deepseek-v4-flash-vision-exp, alongside version 0.1.1 of DeepSeek Harness, the company’s agent framework, which ships with out-of-the-box support for the new model.

What DeepSeek actually shipped

V4-Flash-Vision-Exp takes the text-only V4-Flash — a 284B-parameter mixture-of-experts model with roughly 13B active parameters per token and a one-million-token context window — and bolts on native image understanding. According to DeepSeek’s release notes, the experimental model “matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning, and world knowledge” while making “a major leap over V4-Flash” on multimodal agent benchmarks, bringing performance “close to Opus-4.8.”

The practical details matter for developers:

  • Image billing: images are tokenized for billing at up to 384 tokens each, charged at standard V4-Flash pricing.
  • API surface: the model supports Chat Completions, Messages, and Responses APIs.
  • Input formats: mixed text + image input, with images accepted via base64 encoding, external URLs, or the Files API.
  • Files API: DeepSeek also launched a Files API in tandem — free to use, letting developers upload an image once and reference it by file_id across requests.

That last point is easy to overlook but commercially significant. Reusing images by reference rather than re-uploading cuts request bandwidth dramatically for agent workflows that repeatedly inspect the same screenshots or diagrams — exactly the use case this model targets.

Reading the benchmark table honestly

DeepSeek published eleven benchmark results comparing the new model against Opus 4.8 and its own text-only V4-Flash. The headline: V4-Flash-Vision-Exp beats Opus 4.8 on three of the eleven benchmarks — DeepSWE by 1.3 points, Agents’ Last Exam by 1.6, and ZeroBench by 1.0. On the other eight it trails, sometimes narrowly: Toolathlon-Verified splits 75.9 to 76.2, Chartography 64.3 to 65.0, Terminal Bench 2.1 lands at 83.9 against 85.0.

Two gaps are wide. On NL2Repo, a repository-scale coding task, DeepSeek scores 57.7 against Opus 4.8’s 69.7 — a twelve-point deficit. DSBench-Hard goes 63.6 to 71.7. Those are exactly the kinds of hard, long-horizon tasks enterprises are buying agents to do, and the gap there is real.

There is also an asterisk worth reading. DeepSeek’s own footnote on the multimodal leap explains that on ApexBench and Agents’ Last Exam, the text-only V4-Flash was scored on tests containing images it structurally cannot see — it “ignores multimodal elements contained therein.” The jump from 26.2 to 36.5 on ApexBench is partly a measurement of what happens when you give a blind model an eye test. To DeepSeek’s credit, the company disclosed this in the table rather than leaving it to be discovered — more transparency than many labs offer.

Perhaps the most interesting finding is one DeepSeek undersells. The company says the vision model “matches” V4-Flash on text, but on its own figures the vision variant beats the text-only model on six of seven text benchmarks: Toolathlon-Verified improves by 5.6 points, DeepSWE by 4.9, DSBench-Hard by 4.0. The sole exception is CyberGym, a security benchmark, where adding sight cost 1.4 points (75.3 vs 76.7). Multimodal training, in other words, appears to have transferred positively to text capabilities rather than degrading them.

One benchmark deserves a note for what it says about the entire field: on AutomationBench, all three models score in the mid-twenties — 25.7, 25.1, and 27.2. Whatever today’s agents are good at, that task is not it.

Which Opus, and why the comparison is narrower than it looks

A critical detail in the framing: Anthropic released Claude Opus 5 on July 24, 2026. Opus 4.8, which arrived in May, remains listed as Active on Anthropic’s deprecation page — fully supported and recommended, with no retirement before May 2027 — so DeepSeek picked a current, supported competitor rather than an abandoned one. Five Opus models carry Active status simultaneously.

But DeepSeek’s table contains no Opus 5 column, and no third party has published that comparison either. How V4-Flash-Vision-Exp stacks up against Anthropic’s newest frontier model is simply unknown. The release doesn’t claim otherwise — DeepSeek said “close to Opus-4.8,” and on several benchmarks it is. It did not say “close to Anthropic’s newest model.”

All figures also come from DeepSeek itself, evaluated with its own harness in minimal mode, temperature 1.0, top_p 0.95. The settings are disclosed — vendor benchmarks are normal practice — but they are still the vendor’s.

The commercial argument

The reason this release lands despite the trailing numbers is price. Research earlier in August found V4-Flash to be the cheapest well-known model to run: roughly $0.87 per million words from DeepSeek against approximately $50 from Anthropic — a gap corporate buyers have already noticed. Against a spread like that, trailing by a point or two on most benchmarks is a commercial argument rather than a defeat. Losing NL2Repo by twelve points is a different matter.

DeepSeek separately confirmed that the official V4-Pro GA shipped on August 13 with significantly enhanced agent capabilities, Responses API support, and Codex integration — the model most enterprise buyers would actually weigh, and notably absent from Friday’s table.

Why it matters

The pattern here is now familiar: a Chinese lab ships a capable open-ecosystem model at commodity prices, benchmarks it against a slightly-older-but-still-supported American flagship, and lets the price gap do the talking. V4-Flash-Vision-Exp is experimental — the “Exp” suffix signals exactly that — but it demonstrates that DeepSeek’s multimodal agent story is advancing quickly, and that the intelligence-per-dollar frontier continues to be pushed hardest from Hangzhou. The test that would settle the real question — V4-Flash-Vision-Exp versus Opus 5 on a neutral harness — has a name and no results yet. Until then, an experimental multimodal model sitting within a few points of a supported American frontier model on most benchmarks, at a fraction of the cost, is a genuine market event.