A 560B Model for Free: Ant Group Puts Ling-3.1-flash on OpenCode
Ant Group's InclusionAI made its 560B-parameter MoE Ling-3.1-flash free on OpenCode, ranking No. 2 among open models on Mobile App Arena — with an unresolved context-window gap and an unconfirmed license.
The most interesting free lunch in AI this week is a Chinese one. AntLingAGI, the model team inside Ant Group — the payments giant behind Alipay — has made Ling-3.1-flash, a 560-billion-parameter open-weights model, available free of charge on OpenCode, the open-source coding-agent platform. For anyone running agent loops that burn through tokens, a no-cost model near the open-weights frontier changes the economics of experimentation overnight.
What Ling-3.1-flash actually is
The numbers describe a large mixture-of-experts (MoE) architecture. Ling-3.1-flash has roughly 560 billion total parameters, but only about 25 billion are active per token. Each token routes through a small subset of the model’s experts, which is precisely why a model this large can be hosted cheaply enough to give away: inference cost tracks the 25B active figure, not the 560B total. It is the same sparse-design playbook that DeepSeek, Kimi, and Qwen have used to keep Chinese open models competitive on cost.
The model was first announced on September 30, 2026, and InclusionAI says it is positioned for general-purpose agents, search, routine office work, and software/code development — an agentic and productivity framing rather than a reasoning-benchmark framing. That matches how Ant organizes its model family: Ling is the general-purpose text-and-tool line, Ring is the reasoning line, and Ming handles multimodal.
The scale-up over its predecessor is dramatic. Ling-3.0-flash, shipped in August with MIT-licensed weights, has 124B total and 5.1B active parameters. Ling-3.1-flash is about 4.5× the total parameter count and 4.9× the active count — Ant has roughly preserved the sparsity ratio while scaling the whole body, buying capacity and world knowledge rather than necessarily deeper per-token reasoning.
The scoreboard — and how to read it
Ling-3.1-flash arrives with several vendor-reported scores:
| Benchmark | Reported result |
|---|---|
| Mobile App Arena (Design Arena) | No. 2 open-weights, 1,207 Elo (No. 17 overall) |
| GDPVal-AA v2.1 | 1,673 Elo |
| FrontierSWE | 75.16 |
| HealthBench Professional | 65.35 |
| Terminal-Bench 4.0 | 40.4% |
| SWE-Atlas | 55.9% (part of an 81.0 agent aggregate) |
Three caveats apply. First, these are Ant’s own numbers — no independent lab has reproduced them yet, and there is no Artificial Analysis page for the model. Second, arena rankings are preference-based: the Mobile App Arena specifically measures how people judge generated mobile-app output, so 1,207 Elo is a statement about UI-building polish, not general intelligence. Third, the “No. 2 open” headline has a sober companion: No. 17 overall, a reminder that the closed frontier remains well ahead. For reference, Epoch’s Capabilities Index currently tops out at Claude Opus 5.5 with a score of 167.
One footnote Ant’s own card adds: the HealthBench Professional result was evaluated in what the company calls the “AQ environment,” and no public documentation explains what AQ is. A headline score carrying an unnamed environment qualifier is not something an outside team can reproduce.
The context-window gap you should plan around
Here is the discrepancy that matters in practice. The developer has described capability up to 1 million tokens of context, but the two-week free trial that launched the model serves 256K — and OpenCode’s listing shows 262K. Both things can be true: a model can be trained for longer context than a host chooses to serve. But if your workflow depends on feeding entire large repositories into a single shot, the practical guidance is simple: design for 262K and treat anything longer as unverified until the developer confirms the served window.
There is also an availability nuance worth stating plainly. When Ling-3.1-flash launched, the weights were not published — no Hugging Face repository, no license, nothing to download. Ant’s announcement said only “we plan to open-source the model soon.” The model is described as open-weights, and as of this week you can use it free through hosted channels (OpenCode; OpenRouter and Vercel’s AI Gateway also list a free variant, and Novita carries it at zero price), but the exact license terms and downloadable weights remain the items to verify before you route production work through it. Ling-3.0-flash, by contrast, is fully shipped: MIT weights on Hugging Face since August 7, with BF16 (~255 GB) and FP8 (~128 GB) formats plus community quants.
Why “free on OpenCode” matters more than it sounds
OpenCode is an open-source coding agent, and coding agents are the most token-hungry workload in modern AI usage. A single agent loop can burn tens of thousands of tokens of accumulated tool output before taking one action — which is exactly why a 1M-token window is a functional requirement for agentic work, and why a free, capable model changes the math for experimentation. A five-task trial (bug fix, test-writing pass, refactor, UI component, docs edit) run against your current default model costs nothing but time.
The free-tier fine print deserves a read before you paste in proprietary code, though. OpenCode’s own documentation notes that during Ling 3.1 Flash’s free period, collected data may be used to improve the model. Free tiers also throttle — long agent runs can stall mid-task — and free offers end. Keep your agent configuration portable so you can swap models when the terms change.
The bigger picture
Ling-3.1-flash sits inside a broader pattern: Chinese labs shipping very large, very sparse open models at aggressive prices — or no price at all — as a way to build ecosystem gravity around their tooling. Ant Group’s InclusionAI is a sibling effort to the Qwen family in the same broader Alibaba-adjacent ecosystem, and its release cadence (Ling-2.6-flash in June, Ling-3.0-flash in August, Ling-3.1-flash now) is now measured in weeks, not years.
For teams outside China, the decision framework is unchanged: provenance and data-residency policies apply to hosted free tiers as much as to any other cloud dependency, and the censorship-audit debates that have touched other Chinese open models will presumably reach Ling too once weights are downloadable and testable.
Bottom line: Ling-3.1-flash is a large, sparse, open-weights model that is free to try right now and ranks second among open models on a mobile-app arena. That makes it a zero-cost addition to your evaluation set this week. Verify the license when it lands, plan around 262K of context, keep sensitive code off free tiers — and run the trial before the free window closes.
Sources
- [1] https://x.com/AntLingAGI/status/2105335205741596911
- [2] https://technode.com/2026/09/30/ant-group-launches-ling-3-1-flash-with-560-billion-parameters/
- [3] https://explainx.ai/blog/ling-3-1-flash-open-weights-free-opencode-2026
- [4] https://opencode.ai/docs/zen/
- [5] https://www.orcarouter.ai/blog/ling-3-1-flash-vs-ling-3-0-flash
- [6] https://openrouter.ai/inclusionai/ling-3.1-flash