← All posts / Models

Vision as a Feedback Loop, Not Just an Input: Ant Group Open-Sources the 124B Ling-3.0-flash-VL Under MIT

Ant Group's inclusionAI has open-sourced Ling-3.0-flash-VL, a 124B-parameter multimodal MoE that activates only 5.5B per token, reads images and video over a 256K context, ships in BF16 and FP8, and carries a permissive MIT license.

Vision as a Feedback Loop, Not Just an Input: Ant Group Open-Sources the 124B Ling-3.0-flash-VL Under MIT

Ant Group’s open-model lab inclusionAI (better known as AntLing) has quietly completed a trifecta. Having shipped the Ling-3.0-flash language model in early September and followed it days later with the finance-tuned Ling-3.0-flash-Fin, the team has now open-sourced the piece that ties the family together: Ling-3.0-flash-VL, a natively multimodal model that treats vision as a first-class participant in reasoning rather than a bolted-on input channel.

The release is easy to underestimate on paper — a “flash tier” model, not a frontier flagship. But the details rewards a closer look: 124B total parameters in a sparse Mixture-of-Experts configuration that activates only 5.5B per token, native image and video understanding, a 256K-token context window, day-one serving stacks, FP8 weights, and — most unusually for a model of this scale — a permissive MIT license.

The headline numbers

SpecValue
Total parameters124B (sparse MoE)
Active parameters per token5.5B (~4.4% of capacity)
Context window256K tokens
ModalitiesText, image, video
WeightsBF16 + FP8 (sharded safetensors)
LicenseMIT
ServingSGLang (day-one Docker image), vLLM (custom repo)

The BF16 checkpoint totals roughly 250 GB across shards, and the recommended SGLang recipe serves the full 256K context via YaRN on four 141 GB-class GPUs — H200s or H20-3es — or on 4-GPU Blackwell nodes such as B300/GB300 systems. In other words: this is not a laptop model, but it is a single high-end node model, which is precisely the deployment envelope that makes open multimodal weights practical.

Architecture: built for long multimodal agent histories

Three design choices stand out from the model card.

First, the hybrid backbone. The 42-layer transformer alternates KDA (Kolmogorov–Arnold-inspired attention) layers with Gated Multi-Head Latent Attention (MLA) layers at a 5:1 ratio. The stated goal is efficient long-context processing across text, images, video, and — notably — “extended agent task histories.” The architecture was designed with agentic workloads in mind, not adapted to them after the fact.

Second, VideoRoPE for temporal encoding. Rather than treating video as a bag of frames, the rotary position encoding scheme encodes both spatial positions and temporal order. The card explicitly lists event localization, long-video question answering, and video clip editing as supported tasks — capabilities that usually live in dedicated video models, not general-purpose open weights.

Third, the visual pipeline. A ViT encoder extracts features from images and video; a two-layer MLP projector aligns those features with text representations for unified multimodal understanding. Standard stuff in 2026 — but the framing around it is not.

Understand, Reason, Act — and verify

The most interesting claim in the release is architectural philosophy rather than a benchmark. Ling-3.0-flash-VL is described as integrating vision into “the complete process of understanding, reasoning, planning, acting, and verification.” The team breaks its capabilities into three dimensions:

  • Understand — comprehending complex visual information: object counting, complex layouts, charts, and document content.
  • Reason — reasoning with visual evidence: using what the model sees for calculation, multi-step reasoning, and external information verification.
  • Act — interacting with interfaces: reading web and software UIs and translating them into sequences of actions.

That third dimension is the tell. GUI agents have become one of the most commercially contested categories in AI, and a vision model that closes the loop — perceiving a screen, reasoning about it, acting, then verifying the result visually — is exactly the substrate that category needs. Chinese-language coverage of the release described this as a “visual feedback loop” that reinforces web page generation and GUI task execution. Thinking mode is enabled by default (temperature 0.6, top-p 0.95), consistent with a model intended to deliberate before acting.

What the benchmarks say — and what they don’t

The model card reports a score of 42 on the Artificial Analysis Intelligence Index v4.1.1, a 4-point improvement over Ling-3.0-flash’s 38 — a notable result, since it means adding vision improved the model’s overall intelligence score rather than diluting it, a failure mode that has historically plagued multimodal merges. Artificial Analysis itself flagged the model at launch week with a score of 25 on its then-current index generation; the card’s v4.1.1 figure reflects the newer methodology. Terminal-Bench 2.1 was evaluated under the Artificial Analysis protocol with the Terminus 2 harness, a unified 2-hour timeout, and three runs per task.

Independent commentary has been measured rather than euphoric — which is arguably the right register. An Orcas Router comparison with the text-only Ling-3.0-flash makes the sensible point that the two are complements, not competitors: if your workload never touches pixels, the smaller-footprint base model remains the rational choice; VL’s value appears exactly when perception enters the loop.

The MIT license is the story under the story

At a moment when licensing terms are among the most contested variables in open-weight AI, a 124B multimodal model under MIT is a statement. It places essentially no restrictions on commercial use, modification, or redistribution — compare the custom community licenses attached to many nominally “open” frontier models this year. Combined with FP8 weights (which roughly halve serving memory versus BF16) and a free tier on OpenRouter alongside paid API access on providers like DeepInfra and Novita, the friction to try Ling-3.0-flash-VL is close to zero, whether you self-host or call an API.

Context: the Chinese efficiency wave, item by item

This blog has tracked what might be called the Chinese open-model efficiency thesis — DeepSeek’s V4.1-Flash with its 1M-context KV-cache innovations, Qwen’s relentless cadence, and Ant’s own Ling line. Ling-3.0-flash-VL is a clean, specific data point for that argument: rather than chasing the absolute frontier, the design bets that a 4.4%-active MoE with disciplined architecture can deliver most of the useful capability — here, multimodal agentic perception — at a fraction of the serving cost, and that openness under a permissive license compounds the advantage through ecosystem adoption.

Whether that bet pays off will show up less in leaderboards than in deployment: if Ling-3.0-flash-VL starts appearing in GUI-agent stacks and document-parsing pipelines where GPT-class and Gemini-class models are too expensive to run at volume, the flash tier will have done its job.

For now, the weights are on Hugging Face and ModelScope, the SGLang and vLLM recipes are documented, and the model is free to try on OpenRouter. That is, by any standard, a genuinely open release.