← All posts / Models

DeepSeek Gives Its Cheapest Model Eyes: V4-Flash-Vision-Exp Lands Within Striking Distance of Opus-4.8

DeepSeek's experimental multimodal model adds image understanding to its 284B-parameter MoE at V4-Flash prices, beating Anthropic's Opus-4.8 on three of eleven agent benchmarks and matching it on several more.

DeepSeek Gives Its Cheapest Model Eyes: V4-Flash-Vision-Exp Lands Within Striking Distance of Opus-4.8

For two years, the multimodal frontier has been an expensive neighborhood. Vision-capable agents from Anthropic, OpenAI, and Google sit at the top of the pricing ladder, and the cheapest capable models have been text-only. On August 21, 2026, DeepSeek quietly broke that arrangement: the Hangzhou lab shipped DeepSeek-V4-Flash-Vision-Exp, an experimental model that grafts image understanding onto its bargain-priced V4-Flash backbone — and then published a benchmark table showing it trading blows with Anthropic’s Opus-4.8.

The release is live now on DeepSeek’s API platform under the model name deepseek-v4-flash-vision-exp, alongside version 0.1.1 of the company’s DeepSeek Harness with out-of-the-box support. It is an experimental build, and the “Exp” suffix is doing honest work — but the numbers deserve a careful read, because they say as much about where the whole agent market is heading as they do about any single model.

What shipped

V4-Flash-Vision-Exp takes the text-only V4-Flash — a 284-billion-parameter mixture-of-experts model that activates roughly 13 billion parameters per prompt — and adds the ability to read images and screenshots, then act on what it sees. DeepSeek’s framing is conservative on the text side: the company says the experimental model “matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning, and world knowledge,” while making “a major leap over V4-Flash” on multimodal agent benchmarks, “bringing multimodal agent performance close to Opus-4.8.”

The API integration is notably developer-friendly. Images are tokenized for billing at up to 384 tokens each, priced identically to V4-Flash. The model supports Chat Completions, Messages, and Responses API styles, accepts mixed text-and-image input, and can ingest images as base64, external URLs, or through the Files API — which also went live with this release, free to use, letting developers upload an image once and reference it by file_id across requests.

What the numbers actually say

DeepSeek published eleven benchmark results against Opus-4.8. The vision model wins three: DeepSWE by 1.3 points, Agents’ Last Exam by 1.6, and ZeroBench by 1.0. On the other eight it trails, sometimes narrowly — Toolathlon-Verified splits 75.9 to 76.2, Chartography 64.3 to 65.0, Terminal Bench 2.1 lands at 83.9 against 85.0 — and sometimes widely: NL2Repo has DeepSeek at 57.7 versus 69.7, a 12-point gap, and DSBench-Hard at 63.6 against 71.7.

The multimodal “leap” carries an asterisk that DeepSeek itself supplied. On ApexBench the new model scores 36.5 against V4-Flash’s 26.2; on Agents’ Last Exam, 27.3 against 25.2. But the company’s own footnote notes that in those evaluations the text-only V4-Flash “ignores multimodal elements contained therein” — the older model was being scored on tests containing images it cannot see. The leap is real, but part of it measures what happens when you give a blind model an eye test. To DeepSeek’s credit, the disclosure sits in the table rather than buried, which is more transparency than many labs offer.

Arguably the most interesting line item is one DeepSeek undersold. The company said the vision model “matches” V4-Flash on text; its own figures show it beating the text-only model on six of seven text benchmarks. Toolathlon-Verified improves by 5.6 points, DeepSWE by 4.9, DSBench-Hard by 4.0. The exception is Cybergym, where adding sight cost 1.4 points on a security benchmark. All of these figures are vendor-evaluated — DeepSeek tested its own models on its own harness at temperature 1.0 and top_p 0.95, settings it discloses but controls.

One number belongs to the field, not the contest: on AutomationBench, all three models cluster in the mid-twenties — 25.7, 25.1, 27.2. Whatever today’s agents are good at, end-to-end automation is not it.

Why the comparison matters commercially

The reason this release lands is price. Research this month found V4-Flash to be the cheapest well-known model to run: roughly $0.87 per million input words against approximately $50 for Anthropic’s Opus tier — a gap of more than 50x that corporate buyers have already noticed. Against that spread, trailing by a point or two on Toolathlon is a commercial argument rather than a defeat; a 12-point deficit on NL2Repo, repository-scale coding, is a different matter, since that is precisely the workload enterprises are buying agents to do.

There are honest caveats. Opus-4.8 shipped in May; Anthropic has since released Claude Opus 5, and no published table — DeepSeek’s or anyone else’s — includes an Opus 5 column. DeepSeek also confirmed this week that the general-availability build of V4-Pro has shipped with enhanced agent capabilities, Responses API support, and Codex integration, and that model is absent from Friday’s table too. What the release does establish is a price-performance floor: for pennies per million tokens, developers now get an agent that can see screenshots, parse charts, and operate a terminal at within a couple of points of a frontier model on most agentic tasks.

That is the pattern DeepSeek has repeated since V3 and R1: take a capability the frontier labs price as premium, ship a version within striking distance at a fraction of the cost, and let the market do the rest. V4-Flash-Vision-Exp extends the pattern to multimodal agents. If history rhymes, the “experimental” tag will quietly disappear within a quarter — and the premium tier for vision agents will have to justify itself all over again.

Sources