← Back to all posts
News

Mistral Launches Large 4 Preview: 1T Params, 49B Active, Open Weights Due End of October

October 7, 2026 · 04:11 UTC · News
Mistral Launches Large 4 Preview: 1T Params, 49B Active, Open Weights Due End of October

TL;DR

Mistral opened a public preview of Mistral Large 4 on October 6: a 1-trillion-parameter mixture-of-experts model with 49 billion active parameters, native vision, and a 1M-token context window, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters. The API is live now at a half-price launch rate ($0.68 in, $2.09 out per million tokens against a $1.36/$4.18 list). The weights are promised for the end of the month, the license is unpublished, and the independent indexes put the preview level with DeepSeek V4.1 Flash and a few points behind Kimi K3. Mistral calls it "le Chonk." The name started as a hoax.


What Mistral Actually Shipped

The model page lists the preview as mistral-large-4: 1.05T total parameters, 52B active (49B per token plus embeddings and output layers, per the Hugging Face placeholder), a 1.6B vision encoder, and 1M tokens of context. Mistral describes it as a hybrid instruct-and-reasoning model. Simon Willison found the API exposes exactly two reasoning levels, none and high, and that high produced fewer output tokens than none on his pelican test while drawing a better bicycle.

Sparse MoE at this ratio means the model carries a trillion parameters of knowledge but consults about 5% of them for any given token. Think of a hospital that keeps a thousand specialists on payroll but routes each patient to roughly fifty of them: you pay for the building, not for every doctor on every visit. That is why a 1T model can run at 116 output tokens per second while a dense model that size would crawl.

Pricing is the first thing to read twice. The list price is $1.36 per million input tokens and $4.18 per million output, which is what Artificial Analysis records. The docs page shows those figures struck through next to a preview rate of $0.68 input, $0.07 cached input, and $2.09 output. Budget on the list price; the discount has no stated end date.

It is available today through Mistral Studio, served from the same European cluster it was trained on. Mistral says the model was trained with the same RL environment it sells to customers as Mistral Forge, that training data spanned more than 160 languages including every official EU language, and that until the weights ship it is red-teaming the model with cybersecurity firms, vetted partners, and state authorities, who get a version with reduced moderation and expanded cyber capabilities.

Oct 6, 2026preview API live Octoberpartner red-team end of Octoberweights on HF
The preview is a product; the open-weight release is a date. The Hugging Face placeholder says October 31.

The Benchmarks Mistral Picked

Mistral leads with cybersecurity, and the framing is pointed. On the Artificial Analysis Cyber Index it says ML4 ranks in the global top five and leads every open-weight model built outside China. On the index's vulnerability-reproduction task, where a model must reproduce a real flaw in open-source software and then patch it, ML4 scores 82%, which Mistral calls the highest of any model. It also solves 93% of the 40 challenges in Cybench.

The twist is why 82% is a top score: Mistral says Claude Opus 5.5 and GPT-6 Astra land near zero on the same test because they refuse to do it. That is the sales pitch in one line. Defenders need a model that will prove an exploit is real, and a closed vendor's refusal policy can cut that capability off mid-incident. Mistral is selling the absence of a safety filter as a feature, with open weights and on-prem deployment as the enforcement mechanism for your policy instead of theirs.

Coding is solid rather than spectacular: 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4, for a combined Coding Agent Index of 49.8% that Mistral places ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. In a blind Surge AI human evaluation of coding quality, ML4 scored 3.74 out of 5, second of five models: behind Claude Opus 5 (4.22), ahead of GLM-5.3 (3.60), Kimi K3 (3.59), and GLM-5.2 (3.40).

For agents, Mistral reports 59.9% on AutomationBench, a set of 657 business workflows across Gmail, Google Sheets, Slack, and Salesforce, and 1,393 Elo on AA-Briefcase. On vision it claims 42% on the Dense 200 visual-grounding test against 41% for GPT-6 Astra. Safety numbers: 93.3% attack resistance on Lakera's B3 benchmark, 1.691 out of 2 on KORA, and a cyber refusal rate on JailbreakBench, StrongREJECT, and AgentHarm prompts that Mistral says beats every other open model. Capable of the exploit, says Mistral, but polite about it.

Where the Independent Indexes Put It

Artificial Analysis scores the preview 38 on its Intelligence Index, rank 64 of 225 models, with a blended price of $0.79 per million tokens and 116 output tokens per second. For context, Kimi K3 sits at 44 and DeepSeek V4.1 Flash at 39. Mistral Large 3, released in December 2025, scored 9. Willison's verdict: Mistral is "back to being maybe about 6 months behind the frontier," and this is "certainly not a Fable-class model."

Artificial Analysis Intelligence Index (higher is better) Kimi K344 DeepSeek V4.1 Flash39 Mistral Large 438 Mistral Large 39 (Dec 2025) rank 64 of 225 overall :: blended $0.79/M :: 116 tok/s
From 9 to 38 in ten months. Level with DeepSeek's Flash tier, six points behind Kimi K3.

The Vals Index v2.1, updated the day of launch, is less kind. It weights agentic tasks across finance, coding, legal, and tax by each sector's share of US GDP, and ML4 lands at rank 32 with 48.05% accuracy at $13.78 per test. GLM 5.3 scores 53.51% at $7.25 and Kimi K3 50.30% at $6.38. Gemini 4 Argon leads at 68.90%. The cost gap matters more than the accuracy gap: Artificial Analysis flags ML4 as "very verbose," burning 200M output tokens on its evaluation against an 81M median, so cheap list tokens turn into expensive tasks.

Vals Index v2.1 accuracy, with cost per test (Oct 6, 2026) Gemini 4 Argon68.9% :: $15.68 GLM 5.353.5% :: $7.25 Kimi K350.3% :: $6.38 Mistral Large 448.1% :: $13.78 ML4 rank 32 :: GLM 5.3 rank 16 :: Kimi K3 rank 27
Five points behind GLM 5.3 on accuracy, nearly twice the cost per task. Verbosity is the tax.

A Third-Party Spot Check: Plotly's Analytics Exam

Plotly ran the preview through its data-analytics benchmark: 43 questions against a synthetic 20-million-row wind-turbine dataset spread over 44 tables and 554 columns. ML4 answered 32 of 43 (74%) at $1.94 per full run. Mistral Medium 3.5, from April, managed 25 of 43 (58%) at $20.51 per run. Same vendor, five months, 16 points better at roughly a tenth of the cost. Plotly's read was that the model is "not on the Pareto curve yet" but good enough for real analytics work, with GPT-6.1 Sol at $2.61 and Qwen 3.8 27B at $1.71 per run as the cost neighbors.

Plotly analytics exam: cost per full run, 43 questions (lower is better) Mistral Medium 3.5$20.51 :: 58% correct GPT-6.1 Sol$2.61 Mistral Large 4$1.94 :: 74% correct Qwen 3.8 27B$1.71 20M rows :: 44 tables :: 554 columns :: ~22 s per question
Mistral's own April model to its October model: a tenth of the cost and 16 more points on the same exam.

The Weights Question

Mistral's announcement says "weights drop end of this month." The Hugging Face placeholder lists October 31 as the expected date and already has a waiting list. What it does not list is a license. Mistral Large 3, the 675B-total, 41B-active model from December 2025, shipped under Apache 2.0. ML4's page says nothing either way, and a 1T model under a restrictive license is a very different artifact from one you can fine-tune, distill, and resell.

The timing is not an accident. Reflection AI announced Beam the day before: 501B total, 23B active, pretrained on 23.8 trillion tokens, with weights promised "later in October" under Apache 2.0 and a pitch built on 3 to 4 times less inference compute than GLM-5.2. Two Western labs, two days, two half-trillion-plus MoE announcements, both with the files deferred to the end of the month. The open-weight frontier outside China is, for now, a pair of preview APIs and two calendar entries.

Self-hosting math for when the files land: at 8-bit, 1.05 trillion parameters is roughly a terabyte of weights before the KV cache, so this is a multi-GPU node, not a workstation model. The 49B active parameters help throughput, not memory. If you were hoping Strata or SlotStream would stream it off an SSD, the expert count is on your side but the disk is not.

About the name. Decrypt traces "le Chonk" to "Le Chaton Fat," a June hoax that circulated a fictional multi-trillion-parameter Mistral model with fabricated benchmarks. Mistral read the meme and shipped roughly a thirtieth of it. The Hacker News thread passed 1,600 points and 980 comments within the day, split between people relieved that Europe has a model six months behind the frontier and people unimpressed that Europe has a model six months behind the frontier.

What to Do With It This Week

  • Security teams: test the cyber claim on your own CVEs. The 82% reproduce-and-patch score exists partly because closed models refuse. Run the preview against a handful of flaws you have already triaged and see whether it reproduces them, not whether it talks about them.
  • Price per task, not per token. Vals has ML4 at $13.78 per test against $7.25 for GLM 5.3 despite cheaper list tokens. Measure your own output-token counts before moving a workload.
  • Data residency is the real moat. Trained, served, and operated end to end in Europe under European law. If that is on your compliance checklist, ML4 is now the strongest option that ticks it.
  • Do not architect around the weights yet. No files, no license, and Mistral says the model "continues to improve rapidly as we refine it." The October artifact may not match the October 6 preview.

Caveats

  • This is a preview. Every Mistral-reported number can shift before the weights ship, and Mistral says more architecture detail and post-training methodology are coming later.
  • Most comparisons above are Mistral's own evaluations. Terminal-Bench 4 at 28.3% is the weak spot, and it is not comparable to Beam's Terminal-Bench 2.1 figure, which is a different benchmark version.
  • Artificial Analysis lists the context window as 524K while Mistral's docs say 1M. Test long-context retrieval yourself before relying on the larger figure.
  • The license is unpublished. "Open-weight" is a promise until the Hugging Face repo contains files and a LICENSE.
  • The cyber top score depends on competitors refusing the task, and the refusal-rate comparison that balances it is Mistral's own measurement.

Key Takeaways

  • Mistral Large 4 is live in public preview: 1T total parameters, 49B active, 1.6B vision encoder, 1M context, trained on 3,800 Grace Blackwell GPUs in Mistral's own European datacenters.
  • Pricing is $1.36/$4.18 per million tokens at list, discounted to $0.68/$2.09 during the preview with no stated end date.
  • Artificial Analysis scores it 38, up from Mistral Large 3's 9, level with DeepSeek V4.1 Flash (39) and behind Kimi K3 (44). Vals ranks it 32nd at 48.05% and nearly twice GLM 5.3's cost per task.
  • Mistral's lead claim is cybersecurity: 82% on vulnerability reproduction and patching, a score closed models forfeit by refusing.
  • Weights are due at the end of October with no license published. Reflection's Beam, announced a day earlier, promised Apache 2.0 for its own end-of-month drop.

Sources: Mistral: Introducing Mistral Large 4, Mistral docs: Mistral Large 4 model page, Hugging Face: Mistral-Large-4.0-1T05-A52B (upcoming release), Mistral: Mistral 3 (December 2025), Artificial Analysis: Mistral Large 4, Vals Index v2.1, Plotly: Mistral Large 4 data analytics benchmark, Simon Willison: Le chonk, Reflection AI: Introducing Beam, Decrypt: Mistral drops Le Chonk, Hacker News discussion

AIMistralOpen WeightsLLMMixture of ExpertsBenchmarksCybersecurityEurope
CONSOLE
$