Meta Publishes Six Muse Spark Math Papers; Three Results Were Already Found Elsewhere
TL;DR
Meta published six mathematics papers on October 2 that human researchers wrote with Muse Spark 1.1 and 1.2 in Thinking Mode, using the standard Meta AI chat interface with no custom research scaffold. Meta says five of the six answer previously open questions, spanning probability, PDEs, group theory, optimization and non-associative algebra. The catch is in Meta's own footnotes: for three of those five, other teams (one of them an AI agent) had already published or reported the same answer, two via arXiv papers dated August 2026. The result is real evidence that a consumer chat model can do research-grade math with a human driver, and also a reminder that "open problem" has a shelf life.
What Meta actually released
The announcement, also posted by AI at Meta on X, frames it as the next step after gold-medal competition results: can the model help when nobody has a solution path yet? Each paper carries a "Statement of AI Use," lists the human authors, and was checked by separate reviewing mathematicians named on the blog. Meta's stated goal is "not to mass-produce papers, but to empower researchers."
The six papers:
- Probability: The Strict Threshold for Gaussian Ellipsoid Fitting (Aykut Arslan) proves a sharp phase transition at n ~ d2/4 for fitting random Gaussian points to a centered ellipsoid.
- Differential equations: finite-time blow-up for the mass-critical biharmonic NLS (Leonard Dinh) resolves an open question of Boulenger and Lenzmann that Meta dates to 2015, for dimensions N ≥ 2.
- Group theory: Semiabelian Groups Need Not Be Monomial (Joseph Phillip Brennan, Milana Golich) disproves a 2024 conjecture of M. Kida with SmallGroup(384, 20127). Muse Spark wrote the GAP search program that found it.
- Optimization: tightness of the cycle-based relaxation (Arslan) answers a question from Del Pia and Khajavirad about when a binary polynomial optimization relaxation is exact.
- Arithmetic physics: a string two-point function paper (five authors) extends a connection Yuri Manin envisioned in the 1980s. Meta does not count this one as an open-question answer.
- Non-associative algebra: a three-dimensional counterexample (Andres Barei) disproves a conjecture of García-Martínez and Pérez-Rodríguez on evolution algebras.
The overlap, by the dates
To Meta's credit, the blog lists the competing work itself. The arXiv timestamps make the picture plain.
Ellipsoid fitting: three teams got there in August
Theodor Misiakiewicz and Garrett G. Wen posted a proof of the d2/4 SAT/UNSAT transition on August 10. Sofia de la Cerda, Aaron Potechin, Madhur Tulsiani and Jeff Xu posted their own sharp-threshold proof on August 12. Frederic Koehler and Youngtak Sohn followed on August 27 with a universality result that gives the Gaussian threshold of 1/4 as a special case. Meta's paper, which leaves the exact critical point n/d2 → 1/4 open, appeared October 2.
This was a conjecture of Saunderson, Parrilo, Willsky and coauthors that sat open for years, then fell four times in about seven weeks. Like buses, apparently.
Semiabelian groups: an AI agent beat the AI paper
Meta notes that the Nilradical AI agent reported a different counterexample on September 16. Its result is in the kourovka-lean repository as Kourovka Notebook problem 21.68: a semiabelian group of order 2592 with an irreducible nonmonomial representation of degree 8, backed by a kernel-checked Lean proof. Meta's group is smaller (order 384), but Meta's blog does not mention any Lean formalization for its papers; verification is human peer review.
Evolution algebras: same smallest counterexample size
Xing-Yu Hu and Ran Wen posted a one-parameter family of three-dimensional counterexamples to the same conjecture, dated August 12, noting that three is the minimum dimension where it can fail.
None of this means Muse Spark copied anyone. Independent simultaneous discovery is common when a problem gets ripe, and Meta's researchers may have done the work before August. It does mean "answers a previously open question" is doing more work in the headline than the fine print supports for half the claimed set.
What is genuinely interesting here
No scaffold, just the chat box
The most useful detail for builders is the setup. These researchers used meta.ai in Thinking Mode, not a bespoke prover pipeline, agent swarm or Lean loop. Meta's blog describes a workflow where humans chose the problems, the model proposed approaches, wrote search code and drafted sections, and humans checked and rewrote. That is roughly the workflow you already run with a coding agent, applied to proofs.
Think of it as pair programming where your partner is fast, tireless and occasionally confident about things that are false, so you still read every line. The reviewers named on each paper are the code review.
The counterexample pattern keeps winning
Two of the five claims are disproofs by explicit construction: an order-384 group found by a model-written GAP search, and a three-dimensional algebra. Counterexamples are the easiest AI math wins to trust, because checking one is mechanical. The same pattern shows up in Nilradical's Kourovka work. If you want to point a model at math, search-and-verify problems are where it pays off first.
The uncontested two
The biharmonic NLS blow-up result and the cycle-relaxation tightness result carry no overlap note from Meta. Those, plus the Manin-style extension, are where the claim of new mathematics rests. Neither has had time for broad community vetting yet, so treat them as promising preprints, not settled theorems.
Why it matters if you build with models
- Frontier chat apps now do research-grade assistance. No special access was involved; the same capability is in a consumer product.
- Novelty checking is now the bottleneck. When several labs, academics and autonomous agents attack the same open-problem lists, a literature check before publishing matters more than the proof.
- Formal verification is becoming the differentiator. Nilradical shipped a Lean proof; Meta shipped human review. Expect pressure on AI-math announcements to include machine-checked artifacts.
Key Takeaways
- Meta released six papers co-written with Muse Spark 1.1 and 1.2 via the plain meta.ai chat interface, claiming five answer open questions.
- Meta itself acknowledges independent results for three of the five: ellipsoid fitting (three August arXiv papers), semiabelian groups (Nilradical, September 16) and evolution algebras (Hu and Wen, August 12).
- The biharmonic NLS blow-up and cycle-relaxation results have no overlap noted and are the strongest new claims.
- Counterexample search remains the most reliable AI math win; Muse Spark wrote the GAP search that found the order-384 group.
- Meta's papers rely on human peer review, while the competing Nilradical result ships a kernel-checked Lean proof.
Sources: Meta AI Research blog, AI at Meta on X, Brennan and Golich paper, Arslan ellipsoid paper, Misiakiewicz and Wen (arXiv), de la Cerda, Potechin, Tulsiani and Xu (arXiv), Koehler and Sohn (arXiv), Hu and Wen (arXiv), kourovka-lean (GitHub), Runtime Wire