DeepSeek Open-Sources Its First Multimodal Agent: V4-Flash-Vision-Exp Weights Go Public Under MIT
DeepSeek published the full 305B-parameter DeepSeek-V4-Flash-Vision-Exp checkpoint to Hugging Face under an MIT license — native FP8 weights, a 1M-token context, DSpark speculative decoding, and agent scores within a point of Opus-4.8 on five benchmarks.
Ten days after surprising everyone with a cheap experimental multimodal model on its API, DeepSeek has done the part that actually changes the game: it gave the weights away. On August 31, the Hangzhou lab uploaded the complete DeepSeek-V4-Flash-Vision-Exp checkpoint to Hugging Face under an MIT license — the first open-weights multimodal model in the DeepSeek-V4 family, and one of the very few open models anywhere that can look at a screenshot and then go operate a computer.
By the morning of September 1, the repository had already racked up roughly 18,000 downloads and 420+ likes. For a 305-billion-parameter beast that needs a multi-GPU node to breathe, those are remarkable numbers — and a clear signal of how starved the open-source ecosystem is for capable vision-language agents.
What exactly was released
The repository is more than a pile of safetensors. DeepSeek shipped the full running kit:
- The complete checkpoint — 304.6 billion parameters total, stored natively in FP8 (weights for the routed experts sit in a 4-bit FP4 format, with attention and shared components in FP8/BF16), alongside
config.json, tokenizer files, and a safetensors index. - A minimal PyTorch inference implementation covering the vision encoder and aligner, the DFlash attention kernel path, the MoE routing, Hyper-Connections, and the DSpark forward pass.
- Prompt encoding references that don’t depend on PyTorch at all — OpenAI-style JSON content blocks and the compact
<image>path</image>text notation both encode to identical prompts and token IDs. - First-party recipes for vLLM and SGLang, including a one-command Docker serve on a single 4×GB300 node.
The architecture, per the published config.json, is a 43-layer Mixture-of-Experts model with 256 routed experts (6 active per token plus 1 shared), 64 attention heads with a single KV head (aggressive multi-query attention), MLA-style low-rank Q/K projections at rank 1024, 1,048,576 max position embeddings (a 1M-token context window via YaRN scaling from a 64K base), and a 3-layer DSpark speculative-decoding head — the same “draft and verify from one checkpoint” trick DeepSeek pioneered on the V4 text models, which lets vLLM serve it with 3 speculative tokens and adaptive verification.
The license is the headline for many readers. MIT is about as permissive as it gets: commercial use, redistribution, fine-tuning, distillation, all fine. DeepSeek could have chosen a research-only or community license and nobody would have blinked — several recent Chinese open models did exactly that. It didn’t.
The benchmark table, and why it matters
The model card includes a comparison against its text-only sibling (DeepSeek-V4-Flash-0731) and Anthropic’s Opus-4.8, and the numbers are worth reading closely.
On text agent benchmarks — Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, DSBench-Hard, AutomationBench — Vision-Exp holds its own against the July text model and stays within striking distance of Opus-4.8: it actually beats Opus on DeepSWE (59.3 vs 58.0) and essentially ties it on Toolathlon-Verified (75.9 vs 76.2). Adding eyes did not make it worse at typing.
On multimodal agent benchmarks is where the release earns its name. ApexBench Pass@1 jumps from 26.2 (text-only model, images ignored) to 36.5, closing most of the gap to Opus-4.8’s 39.4. On Agents’ Last Exam, Vision-Exp scores 27.3 — ahead of Opus-4.8’s 25.7. On ZeroBench (Pass@5) it leads Opus 35.0 to 34.0, and on Chartography it’s a near-tie at 64.3 vs 65.0.
In other words: an open-weights, MIT-licensed model is now at or above frontier-closed-model level on half the multimodal agent benchmarks published, and within two points on the rest. That was unthinkable for open vision models a year ago, when “open multimodal” mostly meant OCR-plus-image-captioning quality.
What it takes to actually run it
Here is the catch, and it’s a real one. A 304.6B-parameter checkpoint, even FP8-native, is not a laptop workload. DeepSeek’s own vLLM recipe targets a 4×GB300 node; the SGLang cookbook’s flagship configuration runs on B200 GPUs with FP4 quantization. Back-of-envelope: the FP8 checkpoint’s expert weights alone account for ~296B of the parameters, so the download runs to hundreds of gigabytes before you’ve provisioned a single KV cache entry (mitigated somewhat by FP8 KV cache and block size 256 in the official recipe).
The community moved fast, as it always does for DeepSeek. Unsloth had GGUF quantizations of Vision-Exp up within a day, continuing its pattern from the text-only V4-Flash releases — where the Q8 “lossless” quant fit in 162GB and 3–4-bit builds ran on 110–168GB unified-memory Macs. Even compressed, nobody should expect this on a gaming rig. The realistic deployment story is: rented multi-GPU nodes, an on-prem DGX-style box, or a inference provider — OpenRouter already lists the model — for anyone who wants the data-control story without owning the silicon.
That’s precisely the audience this release serves: enterprises and research groups that need a top-tier vision-language agent they can self-host, inspect, fine-tune, and air-gap. The MIT license makes all four trivially legal.
Why this release lands harder than the API launch
When Vision-Exp appeared on the DeepSeek API on August 21, the story was “cheap competitor to Opus-class multimodal agents.” Nice, but API access is rental — DeepSeek controls the pricing, the availability, and the logging. Weights change the relationship:
- The price floor is now structural. Any provider can serve this model, and competition among them will grind margins toward hardware cost. Anthropic and OpenAI’s vision-agent pricing now has an open alternative breathing down its neck at every procurement conversation.
- Distillation is legal and easy under MIT. Expect a wave of smaller multimodal agent models trained on Vision-Exp outputs over the next quarter, the same way V3/R1 spawned an ecosystem in 2025.
- The agent stack is included. Serving recipes with DSpark speculative decoding, tool-call parsers, and reasoning configs mean the open model isn’t a research artifact — it’s deployable infrastructure on day one.
- It keeps pressure on the open-weights race. GLM-5.3-Flash shipped August 26, Tencent open-sourced the 770B Hy4 preview August 28, and now DeepSeek answers with the first open V4-family multimodal model. Chinese labs are shipping open flagship-tier models on a weekly cadence, and each one resets expectations for what “open” has to mean.
The caveats
It’s called Exp for a reason. The card is explicit that this is an experimental checkpoint — continued training bolted onto V4-Flash rather than a ground-up multimodal build. Text-agent numbers dipped slightly versus the 0731 text-only model on a couple of benchmarks (Cybergym 75.3 vs 76.7, NL2Repo still well behind Opus at 57.7 vs 69.7). And “beats Opus-4.8 on two multimodal benchmarks” is not the same as “matches Opus-4.8 as a product” — Anthropic’s model benefits from harness polish, tooling ecosystems, and reliability work that a raw checkpoint doesn’t ship with.
But none of that dilutes the core fact: as of August 31, anyone with the hardware can download, inspect, and deploy one of the world’s better computer-use vision agents — for free, with no strings. The frontier of open multimodal agents just moved to where the closed frontier was a few months ago. That’s the story, and it’s a big one.