← Back to all posts
News

Meta Publishes Six Muse Spark Math Papers; Three Results Were Already Found Elsewhere

October 4, 2026 · 05:06 UTC · News
Meta Publishes Six Muse Spark Math Papers; Three Results Were Already Found Elsewhere

TL;DR

Meta published six mathematics papers on October 2 that human researchers wrote with Muse Spark 1.1 and 1.2 in Thinking Mode, using the standard Meta AI chat interface with no custom research scaffold. Meta says five of the six answer previously open questions, spanning probability, PDEs, group theory, optimization and non-associative algebra. The catch is in Meta's own footnotes: for three of those five, other teams (one of them an AI agent) had already published or reported the same answer, two via arXiv papers dated August 2026. The result is real evidence that a consumer chat model can do research-grade math with a human driver, and also a reminder that "open problem" has a shelf life.


What Meta actually released

The announcement, also posted by AI at Meta on X, frames it as the next step after gold-medal competition results: can the model help when nobody has a solution path yet? Each paper carries a "Statement of AI Use," lists the human authors, and was checked by separate reviewing mathematicians named on the blog. Meta's stated goal is "not to mass-produce papers, but to empower researchers."

The six papers:

the six Muse Spark papers, by status (per Meta's blog) Ellipsoid fittingopen q :: 3 other proofs exist Biharmonic NLSopen q (2015) :: no overlap noted Semiabelian groupsdisproof :: Nilradical agent too Cycle relaxationopen q :: no overlap noted Evolution algebrasdisproof :: Hu and Wen too String / heightsextension, not an open q copper = independent result acknowledged by Meta
Five claimed open-question answers, three with independent solutions elsewhere: two genuinely uncontested results remain.

The overlap, by the dates

To Meta's credit, the blog lists the competing work itself. The arXiv timestamps make the picture plain.

Ellipsoid fitting: three teams got there in August

Theodor Misiakiewicz and Garrett G. Wen posted a proof of the d2/4 SAT/UNSAT transition on August 10. Sofia de la Cerda, Aaron Potechin, Madhur Tulsiani and Jeff Xu posted their own sharp-threshold proof on August 12. Frederic Koehler and Youngtak Sohn followed on August 27 with a universality result that gives the Gaussian threshold of 1/4 as a special case. Meta's paper, which leaves the exact critical point n/d2 → 1/4 open, appeared October 2.

This was a conjecture of Saunderson, Parrilo, Willsky and coauthors that sat open for years, then fell four times in about seven weeks. Like buses, apparently.

Semiabelian groups: an AI agent beat the AI paper

Meta notes that the Nilradical AI agent reported a different counterexample on September 16. Its result is in the kourovka-lean repository as Kourovka Notebook problem 21.68: a semiabelian group of order 2592 with an irreducible nonmonomial representation of degree 8, backed by a kernel-checked Lean proof. Meta's group is smaller (order 384), but Meta's blog does not mention any Lean formalization for its papers; verification is human peer review.

Evolution algebras: same smallest counterexample size

Xing-Yu Hu and Ran Wen posted a one-parameter family of three-dimensional counterexamples to the same conjecture, dated August 12, noting that three is the minimum dimension where it can fail.

2026 timeline: independent results vs Meta's release Aug 10M+W Aug 12dlC+ / Hu+Wen Aug 27Koehler+Sohn Sep 16Nilradical Oct 2Meta: 6 papers arXiv dates for papers; Nilradical date as reported by Meta
Every overlapping result was public before Meta's papers went up.

None of this means Muse Spark copied anyone. Independent simultaneous discovery is common when a problem gets ripe, and Meta's researchers may have done the work before August. It does mean "answers a previously open question" is doing more work in the headline than the fine print supports for half the claimed set.

What is genuinely interesting here

No scaffold, just the chat box

The most useful detail for builders is the setup. These researchers used meta.ai in Thinking Mode, not a bespoke prover pipeline, agent swarm or Lean loop. Meta's blog describes a workflow where humans chose the problems, the model proposed approaches, wrote search code and drafted sections, and humans checked and rewrote. That is roughly the workflow you already run with a coding agent, applied to proofs.

Think of it as pair programming where your partner is fast, tireless and occasionally confident about things that are false, so you still read every line. The reviewers named on each paper are the code review.

The counterexample pattern keeps winning

Two of the five claims are disproofs by explicit construction: an order-384 group found by a model-written GAP search, and a three-dimensional algebra. Counterexamples are the easiest AI math wins to trust, because checking one is mechanical. The same pattern shows up in Nilradical's Kourovka work. If you want to point a model at math, search-and-verify problems are where it pays off first.

The uncontested two

The biharmonic NLS blow-up result and the cycle-relaxation tightness result carry no overlap note from Meta. Those, plus the Manin-style extension, are where the claim of new mathematics rests. Neither has had time for broad community vetting yet, so treat them as promising preprints, not settled theorems.

Why it matters if you build with models

  • Frontier chat apps now do research-grade assistance. No special access was involved; the same capability is in a consumer product.
  • Novelty checking is now the bottleneck. When several labs, academics and autonomous agents attack the same open-problem lists, a literature check before publishing matters more than the proof.
  • Formal verification is becoming the differentiator. Nilradical shipped a Lean proof; Meta shipped human review. Expect pressure on AI-math announcements to include machine-checked artifacts.

Key Takeaways

  • Meta released six papers co-written with Muse Spark 1.1 and 1.2 via the plain meta.ai chat interface, claiming five answer open questions.
  • Meta itself acknowledges independent results for three of the five: ellipsoid fitting (three August arXiv papers), semiabelian groups (Nilradical, September 16) and evolution algebras (Hu and Wen, August 12).
  • The biharmonic NLS blow-up and cycle-relaxation results have no overlap noted and are the strongest new claims.
  • Counterexample search remains the most reliable AI math win; Muse Spark wrote the GAP search that found the order-384 group.
  • Meta's papers rely on human peer review, while the competing Nilradical result ships a kernel-checked Lean proof.

Sources: Meta AI Research blog, AI at Meta on X, Brennan and Golich paper, Arslan ellipsoid paper, Misiakiewicz and Wen (arXiv), de la Cerda, Potechin, Tulsiani and Xu (arXiv), Koehler and Sohn (arXiv), Hu and Wen (arXiv), kourovka-lean (GitHub), Runtime Wire

AIMetaMuse SparkMathematicsAI for ScienceResearchLean
CONSOLE
$