No Model Card, No Price, No Announcement: MiniMax Slips M3.1-Flash-Preview Into MiniMax Code
MiniMax quietly shipped M3.1-Flash-Preview inside MiniMax Code with five reasoning tiers and a gated API, betting the model speaks for itself in China's coding-model price war.
MiniMax just shipped a new frontier-tier coding model, and almost nobody was supposed to notice. On September 27, the Shanghai-based lab confirmed that M3.1-Flash-Preview had gone live inside MiniMax Code, its coding assistant product — a full day after developers first spotted the unannounced model sitting in the product’s model picker. There was no launch event, no blog post with benchmark tables, no model card, no published price per million tokens, and no public API identifier. In an industry that treats every model drop as a press-cycle event, the silence was the story.
What actually shipped
For now, M3.1-Flash-Preview is available only through MiniMax Code and the company’s Token Plan subscription. The API endpoint is gated — the only way to try the model is inside MiniMax’s own product. According to MiniMax’s own API documentation, the model is described as a “frontier multimodal coding model with 1M context window and tunable thinking depth,” and it sits in the picker alongside the older M3, M2.7, and M2.7-highspeed options.
The most technically interesting change is a new reasoning_effort field with five tiers: low, medium, high, xhigh, and max. M3 previously offered a simpler thinking toggle. The expanded control lets developers trade response latency against deliberation on every request — lower tiers for quick edits and routine tool calls, higher tiers for planning, debugging, and repository-wide changes. MiniMax has not documented the compute budget, latency profile, or quality delta for each tier, so teams will have to measure the settings against their own workloads.
The launch also came bundled with a promotional event: from September 28 through October 7 (UTC+8), MiniMax Code users get double free credits for daily check-ins, and all Token Plan quotas were reset the moment M3.1-Flash-Preview went live, with more resets promised as “surprises” during the event window.
What the leaks suggest
A document circulating as a partner architecture note — unverified, and not commented on by MiniMax — outlines several changes from M3. If it reflects the deployed model, the design concentrates squarely on cutting memory traffic and accelerating token generation:
- All-sparse attention across the context window
- 4-bit quantization for experts and the KV cache
- DSpark, a replacement for the EAGLE-style speculative decoder
- A gated, roughly 250 GB Hugging Face repository named
MiniMax-M3.1-preview-private
Those are exactly the techniques a lab would reach for to build a “Flash” tier: cheaper serving, faster interactive latency, and a model cheap enough to throw at dozens of small coding tasks a day.
The contrast with M3
The quiet rollout is a sharp break from how MiniMax shipped its last flagship. When M3 launched in June, the company published architecture details, SWE-bench scores, and a pricing sheet that undercut rivals by as much as 90%. M3 is a 428-billion-parameter mixture-of-experts model with roughly 23 billion active parameters per token, a one-million-token context window, and native image and video input. MiniMax reported 80.5% on SWE-bench Verified and 59.0% on the harder SWE-bench Pro, numbers the company says put it ahead of GPT-5.5 and Gemini 3.1 Pro on that test — though these are MiniMax’s own measurements on its own infrastructure, so they should be read as claims rather than independent verdicts.
On the efficiency front, M3 uses one-twentieth of its predecessor’s per-token compute at a one-million-token context, with more than ninefold faster prompt processing and more than fifteenfold faster decoding. But that scale still imposes substantial serving demands for short, iterative coding tasks — precisely the workload M3.1-Flash-Preview appears built to absorb.
Why ship a model in silence?
Reading the strategy, MiniMax is effectively A/B testing a product feature in public and letting the internet find out. That’s either confidence that the model will speak for itself once developers use it, or a signal that this release isn’t the flagship — just a faster, cheaper sibling of M3 built to eat routine coding work while the real frontier effort happens elsewhere.
Splitting a model line into a large frontier version and a fast, cheap version is not new; OpenAI and Google both do it. What’s notable is that MiniMax is doing it around an open-weight flagship while keeping the faster variant locked inside its own coding tool for now. The product funnel is the point: get developers living inside MiniMax Code on Token Plan subscriptions, where per-token comparisons matter less than they do on open API markets.
The China context
MiniMax isn’t shipping into a quiet market. Alibaba released Qwen3.8-Max on August 3 — a 2.4-trillion-parameter flagship priced at $2 per million input tokens — and followed it on August 12 with the first open-weight version of a Max-tier Qwen model. Zhipu’s GLM-5.3 landed August 14 with a million-token context window of its own. DeepSeek and Kimi are running the same playbook: ship often, price low, publish weights.
And developers are noticing. According to CNBC’s reporting on OpenRouter data, Chinese models accounted for 57% to 67% of total token usage on the platform for the week including September 14 — up from just 6% to 13% in February. Vercel saw a similar jump, with Chinese models’ share of usage rising to 55% in August from 11% in January. The reason is not mysterious: for the agentic coding and customer-service workloads that now consume most enterprise AI budgets, Chinese open models run 60% to 90% cheaper than the leading U.S. alternatives.
Washington is watching that shift with more alarm than enthusiasm. Daniel Remler, a senior fellow in the technology and national security program at the Center for a New American Security, told CNBC that Chinese AI represents “real economic and security risks for the United States” and warned that integrating Chinese models could pull countries into China’s technology sphere of influence.
What’s still missing
The evidence gap remains the story’s asterisk. MiniMax has not published a benchmark report, a model card, pricing, or API availability for M3.1-Flash-Preview. The leaked architecture notes are unverified. Tier labels don’t reveal token budgets or guarantee consistent behavior across tasks. Until the company publishes measurements, developers evaluating the preview should instrument their own — latency, completion quality, tool-call accuracy, and failure rates at each reasoning level — and treat the default max setting with caution, since it may obscure the very speed benefits the Flash name implies.
For production adoption, the missing details matter more than the hype. Systems that require stable identifiers, documented behavior, and support commitments should keep M3.1-Flash-Preview at arm’s length for now. But as a signal of where the coding-model market is heading — Chinese labs shipping faster, cheaper variants at a near-constant cadence, and occasionally not even bothering to announce them — this quiet drop says more than a launch event would.
Sources
- [1] https://alphasignal.ai/news/minimax-quietly-slips-m3-1-flash-preview-into-its-coding-tool
- [2] https://startupfortune.com/minimax-quietly-ships-a-coding-only-model-as-chinas-ai-models-flood-the-market/
- [3] https://promptblueprints.tech/ai-releases/minimax-m3-1-flash-preview-arrives-on-minimax-code/
- [4] https://platform.minimax.io/docs/guides/models-intro