← All posts / Research

Six Papers, Five Open Problems: Meta's Muse Spark Joins the Mathematicians

Meta AI Research published six mathematics papers co-authored with Muse Spark 1.1 and 1.2 in Thinking Mode — five answer previously open problems, from a sharp ellipsoid-fitting threshold to a 384-element counterexample in group theory, all produced through the ordinary meta.ai chat box.

Six Papers, Five Open Problems: Meta's Muse Spark Joins the Mathematicians

For years the benchmark of AI mathematical ability has been the Olympiad: gold-medal scores on problems written for teenagers, where the trick is that somebody already knows the answer. On October 2, Meta AI Research moved the goalposts. It published six research papers co-authored by working mathematicians and its Muse Spark models — and five of them answer questions that, until now, nobody could answer.

The result that will be studied hardest is not on any leaderboard. It is the workflow: the mathematicians used Muse Spark 1.1 and 1.2 in Thinking Mode through the regular meta.ai chat interface, with no custom research scaffold — the same chat box available to any of Meta’s billions of users. No bespoke proof assistant, no cluster of agents, no special tooling. Frontier reasoning, applied to genuinely open mathematics, from a consumer chat window.

From answer keys to no answer key

Meta’s own framing is careful. Earlier this year its models achieved gold-medal-level performance across five high-school Olympiad competitions in mathematics, physics, and chemistry. “Competition problems can be incredibly difficult, but those problems already have a solution,” the company writes. “Open research is different. There is no answer key, no guarantee that an approach will work.”

That difference defines the collaboration’s rules of engagement, and they deserve attention because they are the real template here:

  • A team of mathematicians guided the research and worked with Muse Spark to explore ideas and develop the arguments.
  • A second, independent group of mathematicians reviewed each result.
  • Every paper explicitly marks which passages were primarily drafted by researchers and which by the AI.
  • Each paper credits the earlier mathematical work it builds on — including, notably, competing results arrived at independently.

That last point is not boilerplate. After completing the work, Meta learned that teams outside the company had independently announced solutions to some of the same problems using different approaches. Rather than bury that, the papers acknowledge it directly — a standard of transparency more established labs routinely fail to meet.

The five solved problems

Probability — The Strict Threshold for Gaussian Ellipsoid Fitting. How many random Gaussian points in high dimensions can you fit to an ellipsoid centered at a fixed point? Aykut Arslan, working with Muse Spark, identified a sharp threshold: below it, the ellipsoid exists with high probability; above it, it almost certainly does not. The behavior exactly at the threshold remains open. Three independent groups — Misiakiewicz and Wen; De la Cerda, Potechin, Tulsiani, and Xu; and Koehler and Sohn — posted concurrent results in August 2026 using different techniques.

Differential equations — Finite-time blow-up for the biharmonic nonlinear Schrödinger equation. Leonard Dinh settled a question left open since 2015 about wave collapse in a model inspired by laser physics: for radial solutions with negative energy in two or more dimensions, collapse must occur in finite time. The proof also confirms a prediction from computer simulations dating to 2002. Muse Spark worked through calculations, tested candidate arguments, and helped revise the proof.

Group theory — Semiabelian groups need not be monomial. Kida’s 2024 conjecture said every finite semiabelian group is also monomial. Muse Spark wrote a search program in GAP, the computational algebra system, that found a 384-element counterexample — a black swan for the conjecture. Milana Golich and collaborator Joseph Phillip Brennan verified the result and completed the argument. Meta also credits the AI agent Nilradical, which reported a different counterexample to the same conjecture on September 16, 2026.

Optimization — Tightness of the cycle-based relaxation. A question first posed by Del Pia and Khajavirad in 2026: when does a relaxation of a binary polynomial optimization problem capture the original exactly? The answer, developed with the model’s help in reframing the problem probabilistically, is a clean structural rule — exact when each pairwise-only region contains exactly one decision, and lossy otherwise.

Non-associative algebra — Solvable evolution algebras. The García-Martínez–Pérez-Rodríguez conjecture proposed a test for identifying “solvable” evolution algebras, structures inspired by evolutionary biology. Muse Spark generated a small three-dimensional counterexample that passes the test without belonging to the class — and then helped establish an alternative, subspace-based characterization that replaces the broken rule. Independent counterexamples by Hu and Wen are acknowledged.

A sixth paper, in arithmetic physics, connects the string two-point function to a height function on curves — a bridge between number theory and p-adic string theory following a direction Yuri Manin envisioned in the 1980s. Muse Spark drafted three core technical sections, which the five authors then checked, corrected, and refined.

Why it matters

Three things distinguish this announcement from the steady drumbeat of “AI solves math” claims.

First, the division of labor is documented. These are not papers where an AI is a black box emitting proofs. In the group-theory case, the model’s contribution was concrete and checkable: it wrote the search program that found the counterexample. In the physics case, it drafted sections that humans then repaired. This is the collaboration model the field has argued about for two years, operating at a scale and openness nobody has shown before.

Second, the tooling is ordinary. DeepMind’s AlphaProof and AlphaGeometry stack, and OpenAI’s experimental reasoning setups, are specialized systems. Muse Spark did this work through meta.ai, in Thinking Mode, the way a student would use it for homework. If frontier models in a consumer chat box can materially contribute to open mathematics, the population of people who can attempt serious research just expanded by orders of magnitude.

Third, the timing is not incidental. The papers land the same week former OpenAI safety lead David Robinson published his resignation essay warning the industry is “not being nearly careful enough,” and amid a White House voluntary accord premised on the argument that the sector can police itself. Meta’s choice to publish papers with AI passages explicitly marked, independent human review built in, and rivals’ concurrent results credited is — intentionally or not — a demonstration of what responsible acceleration can look like.

The caveats

The announcement is not a claim that AI has replaced mathematicians. Every paper was guided by researchers who chose the problems and key proof ideas; every argument was verified by humans outside the original team. The behavior at the ellipsoid-fitting threshold remains unresolved. And concurrent independent solutions to two of the problems suggest the field as a whole is moving, not that one model is uniquely responsible.

But the direction is unambiguous. A year ago, AI co-authored mathematics meant benchmark scores. Now it means citable answers to questions that were open last month — produced in a chat window, with the receipts attached.