← All posts / Models

124B Parameters, Open Weights: Ant Group Open-Sources Ling-3.0-flash-Fin and the FinFIRST Benchmark

Ant Group releases Ling-3.0-flash-Fin, a finance-tuned MoE model with 124B total / 5.1B active parameters plus the expert-built FinFIRST benchmark — betting that Wall Street's next research analyst runs on open weights with auditable evidence chains.

124B Parameters, Open Weights: Ant Group Open-Sources Ling-3.0-flash-Fin and the FinFIRST Benchmark

While Western frontier labs chase general reasoning trophies, Ant Group’s AI team has quietly shipped something aimed squarely at one of the most demanding enterprise domains in the world: sell-side and buy-side financial research. This week the company open-sourced Ling-3.0-flash-Fin, the first finance-enhanced model in its Ling family, together with FinFIRST, an expert-built benchmark for financial search agents. Both the model weights and the evaluation harness are now publicly available — a combination that says a great deal about where Ant thinks the next phase of applied AI competition actually happens.

What was released

Ling-3.0-flash-Fin is built on Ling-3.0-flash, Ant’s hybrid-reasoning Mixture-of-Experts base model. The headline numbers: 124 billion total parameters, with only 5.1 billion activated per token. That sparse-activation profile is the whole point. You get the knowledge breadth of a very large model, but inference costs and deployment footprints closer to something far smaller — a trade-off that matters enormously for banks running thousands of concurrent research-agent sessions against private data.

The model was co-developed with financial institutions and domain experts who shaped task design, data systems, and evaluation methods, and it targets four pillars of real financial work:

  1. Information retrieval — prioritizing official and authoritative sources so every data point can be traced from origin to output.
  2. Research reasoning — combining information across multiple sources and formats into verifiable evidence chains rather than confident single-shot answers.
  3. Valuation modelling — navigating complex Excel link chains, supporting automated updates, and keeping files editable rather than freezing them into dead PDFs.
  4. Report generation — assembling facts, calculations, and charts into professional research documents.

Crucially, the release includes open weights. Firms can deploy it privately, behind their own firewalls, and wire it to external tools — search, Python, databases, spreadsheets — instead of being locked to a hosted API. For compliance-heavy institutions, that distinction is not a nice-to-have; it is often the difference between a deployable system and a policy violation.

The benchmark is the story

As significant as the model itself is FinFIRST, open-sourced alongside it and developed with support from the investment banking team at China International Capital Corporation (CICC). FinFIRST V1 ships with:

  • 123 expert-authored tasks drawn from realistic research workflows
  • 701 atomic criteria for granular scoring
  • 12,300 rubric points across the task set

What makes FinFIRST interesting is not the volume but the philosophy. It evaluates the full research process, not just whether a final answer matches a reference output. Where the number came from, how the calculation was constructed, whether each intermediate step can be audited later — these are first-class scoring dimensions, not afterthoughts. As one commentary on the release put it, FinFIRST “turns ‘show your work’ into an actual benchmark dimension.”

This addresses a genuine gap. Financial work sits in the small subset of professional tasks where process is the product. A valuation that arrives at the correct terminal multiple through fabricated sourcing is worse than useless — it is a liability. Existing general-purpose benchmarks, which overwhelmingly score final answers, structurally cannot detect that failure mode. An expert-built rubric with thousands of atomic criteria can.

Competitive positioning

On the numbers Ant reports, Ling-3.0-flash-Fin posts competitive results across a wide bench: FinFIRST itself, FinSearchComp Verified, FinCRAFT, FinanceAgent v1.1/v2, APEX-Agents, SpreadsheetBench v1/v2, and τ3-Banking. The spread is telling — search, spreadsheet manipulation, multi-step reasoning, and agent-based task completion, which is a fair approximation of what a junior analyst actually does all day.

The release also slots into Ant’s broader Ling 3.0 portfolio strategy: a base flash model, a tiny variant for local deployment, a VL variant for visual and document understanding, and now domain-specific forks for finance (Fin) and healthcare (Santé). It is a deliberate horizontal-and-vertical grid — general capability layers underneath, sector-tuned specialists on top, each with benchmarks matched to how that sector actually judges quality.

Why open weights for finance?

The strategic logic deserves attention. Open-weight finance models invert the usual enterprise AI procurement model. Instead of sending proprietary research queries to a third-party API, an institution runs the model inside its own perimeter, keeps every log, and owns the full audit trail. In a post-MiFID-II, post-every-financial-scandal world, auditability is a moat.

There is also a competitive subtext. Western labs have largely moved finance-grade models behind paid APIs with usage policies, and some have restricted finetuning on regulated workflows. Ant releasing a 124B-class finance model with open weights — usable via OpenRouter and Vercel today, weights on Hugging Face and ModelScope — gives every fintech, boutique research shop, and regional bank a credible path to building agentic research tooling without negotiating a frontier-lab enterprise contract. Whether by design or side effect, it also plants a Chinese-built model family at the infrastructure layer of global financial AI.

What to watch

The open question is adoption gravity. Benchmarks published by the same lab that trained the model always deserve skepticism — FinFIRST’s real test will be whether independent institutions treat it as a neutral yardstick or as Ant’s home-field advantage. Watch for third-party FinFIRST leaderboards, and for whether CICC’s involvement pulls other Chinese brokerages into building on the Ling stack.

For everyone else, the signal is simpler: domain-specialized open-weight models with process-level evaluation are now shipping at frontier-adjacent scale. The era of “one general model fits all verticals” is being quietly contested at the layer where it actually gets used — and the first shot in finance just came from Hangzhou, not Silicon Valley.