← All posts / Models

Europe's New Frontier Contender: Multiverse Computing's Quasar 438B Scores 43 on the AA Intelligence Index

The Spanish quantum-inspired AI lab's first large model beats Mistral Medium 3.5 and NVIDIA Nemotron 3 Ultra on Artificial Analysis benchmarks while answering 500-token prompts in 15.3 seconds — Europe's highest-scoring model yet.

Europe's New Frontier Contender: Multiverse Computing's Quasar 438B Scores 43 on the AA Intelligence Index

Europe has a new candidate for its most capable AI model. On September 2, 2026, Multiverse Computing — the San Sebastián, Spain-based company best known for its quantum-inspired tensor-network model compression — released Quasar 438B, its first large language model. It is a 438-billion-parameter reasoning model built for enterprise-scale agents and coding, and according to the company’s own launch materials, it scores 43 on the Artificial Analysis Intelligence Index v4.1.1: the highest result of any European model in the comparison field.

The launch is a statement of intent from a company that until now has made its name on making other people’s models smaller, faster, and cheaper. With Quasar, Multiverse Computing is stepping into the 400B-plus parameter class with its own flagship — and bringing its efficiency-first engineering philosophy along with it.

What the benchmarks actually say

The Artificial Analysis Intelligence Index v4.1.1 is a weighted composite of nine evaluations across four categories: agents (GDPval-AA v2 and τ³-Banking), coding (Terminal-Bench v2.1 and SciCode), scientific reasoning (Humanity’s Last Exam, GPQA Diamond, and CritPt), and general knowledge plus long-context reasoning (AA-Omniscience and AA-LCR). Quasar’s composite score of 43 puts it ahead of Mistral Medium 3.5 at 30, NVIDIA Nemotron 3 Ultra at 38, and Inkling at 42 — a field still led at the top by Claude Opus 5 at 63.

But raw index scores are only half of Multiverse’s pitch. The other half is latency. Quasar returns a 500-token response — thinking time included — in 15.3 seconds. Only three models in the company’s comparison are faster: Nemotron 3.5 Lightning at 9.4 seconds (index score 24), Gemini 3.5 Flash-Lite at 10.8 seconds (37), and Gemini 3.7 Flash at 11.5 seconds (56). Of those, only Gemini 3.7 Flash is both faster and more capable. Of the models that outscore Quasar, only two answer in under 25 seconds; the rest take between 38 and 156 seconds.

The contrast with its European peers is stark. Mistral Medium 3.5 scores 30 and takes 18.8 seconds. Nemotron 3 Ultra scores 36 and takes 25.7 seconds — and carries 112 billion more parameters than Quasar. Inkling scores 42, just one point behind, but needs 48.3 seconds, more than triple Quasar’s response time. Against Mistral specifically, the comparison runs one way on both axes: Quasar scores higher and answers faster.

Long-context reasoning is where Quasar comes closest to the frontier group. It scores 75.0 on AA-LCR, which tests a model’s ability to extract, connect, and reason over information distributed across long documents. That is level with Grok 4.6 (high) at 75.0, within a point of Claude Opus 5 at 75.7 and Qwen3.8 2.4T A95B at 75.3, and comfortably ahead of Nemotron 3 Ultra by 4.0 points and Mistral Medium 3.5 by 9.7 points.

On Terminal-Bench v2.1, which puts agents to work in real terminal environments — inspecting repositories, running commands, diagnosing errors, completing connected sequences of actions — Quasar scores 69.3. That leads Mistral Medium 3.5 by 18.7 points and Nemotron 3 Ultra by 15.4, though it still trails the frontier result led by Claude Opus 5 at 89.1. Multiverse itself flags this as the evaluation with the most headroom and the focus of the next round of work. The company also reports Quasar outperforms comparable European models on eight of the nine individual evaluations that make up the index (seven of eight in its chart selection).

The efficiency DNA behind the model

Multiverse Computing was not founded to train frontier models. The company, led by co-founder and CEO Enrique Lizaso, built its business on CompactifAI — a compression platform that uses quantum-inspired tensor-network techniques to shrink large language models so they can run on edge devices and cheaper infrastructure. The CompactifAI App launched in March 2026 to bring offline AI to edge hardware, and the company’s API offers a curated catalog of compressed, frontier-class models.

Quasar extends that heritage into the flagship class. “This is a significant milestone for European AI: Quasar shows that European AI developers do not have to choose between reasoning performance and speed,” Lizaso said in the launch announcement. “European enterprises need models that can work through complex tasks, use tools and handle long documents, and they also need greater choice and access to powerful AI developed here in Europe.”

That positioning matters commercially. In agentic systems, a single user request may require a model to plan a task, call several tools, check results, and adjust its approach. When an agent makes dozens of model calls to complete one piece of work, per-call delays compound into minutes of additional waiting. A model that holds near-frontier reasoning while keeping response loops fast is exactly what interactive enterprise products need — and exactly what most 400B-class models have historically failed to deliver.

Sovereign AI context

The launch lands in the middle of Europe’s intensifying sovereign-AI push. The AI Insider coverage frames Quasar as evidence that Europe “can compete with leading models from the US and China on capability and speed” — a claim the benchmark table partially supports at the regional level, even as the absolute frontier remains American.

Quasar runs in English and Spanish, making it a practical foundation for European and international enterprises that need one reasoning system across teams and markets rather than a single-language deployment. Multiverse lists software engineering agents, technical copilots, operational automation, document-heavy research, and knowledge work as target applications — domains where accuracy, context retention, and completion time jointly decide whether an AI system is useful in practice.

Access is through the CompactifAI API, listed at $0.60 per million input tokens, so teams can evaluate the model without standing up their own infrastructure. License terms for self-hosting were not stated in the launch note, and Multiverse says more updates for Quasar’s coding and agentic capabilities are planned.

What to watch

Three open questions will determine whether Quasar is a one-off press release or the start of a serious European frontier effort. First, independent verification: these numbers come from Multiverse’s own launch materials, and third-party Artificial Analysis listings will tell whether the 43 holds. Second, the Terminal-Bench gap: 69.3 is strong regionally but 20 points behind Claude Opus 5, and coding agents are the highest-value enterprise workload right now. Third, openness: whether Quasar’s weights, or at least a compressed CompactifAI variant, become available for self-hosting — the move that would actually shift Europe’s sovereign-AI calculus rather than just its leaderboard.

For now, the headline is simple enough: Europe’s efficiency specialist has entered the big-model race, and its first swing lands closer to the frontier than anything the continent has previously shipped.