Mistral AI has released the biggest model in its history. Mistral Large 4, announced on October 6, 2026, is a 1-trillion-parameter mixture-of-experts model that the Paris-based lab calls the most powerful open-weight AI system outside China. It is available now as an API preview, with downloadable weights scheduled for October 27, and it arrives with an unusual emphasis: cyber defence.

The Specs

Large 4, nicknamed "le Chonk" internally, activates about 49 billion parameters per token out of its 1 trillion total, the sparse design that lets very large MoE models run at a fraction of the cost of dense equivalents. Key details reported at launch:

  • Multimodal by design: the model processes text and images natively.
  • 160+ languages: training covered every official European Union language.
  • Training run: roughly two months on about 4,000 Nvidia Grace Blackwell GPUs in Mistral's own European data centres, drawing around 10 megawatts, according to The Next Web.
  • Pricing: API access is listed at $1.36 per million input tokens and $4.18 per million output tokens, according to VKTR.
  • Distribution: besides Mistral's own API, the preview has been added to Vercel's AI Gateway.

It is the first major milestone funded by Mistral's €3 billion Series D round, raised in September.

Benchmarks: Strong, With Caveats

Preliminary results published at launch put Large 4 ahead of leading Chinese open models on several tests. The Next Web reports 62% on DeepSWE v1.1 coding, narrowly ahead of GLM-5.3 at 61% and DeepSeek-V4-Pro at 57%, and 67% on FinWorkBench, tied with DeepSeek-V4-Pro. On Harvey's legal agent benchmark it scored 15%, ahead of Kimi K3, and it led OpenAI's GPT-6 Astra on the DIOR-RSVG visual-grounding test.

The headline strength is security. VKTR reports Large 4 ranks in the global top five on the Artificial Analysis Cyber Index, scored 82% on a test that requires reproducing and then patching a real open-source vulnerability, the highest of any model, and solved 93% of the 40 capture-the-flag challenges in Cybench.

Not every account is uniformly glowing. Coverage citing CNBC says Mistral acknowledged the model still trails closed frontier systems in some areas, including coding. Readers should treat all launch-day figures as vendor-reported until independent evaluations land.

Why the Weights Are Delayed

Mistral is using the three-week window before the open-weight release to red-team a less restricted variant with stronger cyber features alongside security firms, vetted partners and government agencies. Co-founder Guillaume Lample framed the goal as letting enterprises and governments defend themselves against attackers who jailbreak closed models to run cyber attacks.

That staged approach mirrors a broader industry trend. Google's Gemini 4 Argon and OpenAI's GPT-6 Astra were also positioned around cybersecurity in recent weeks, and labs are increasingly gating the most offensive-capable features behind trusted-access programmes. Mistral's twist is that the full model will ultimately be downloadable, which raises both the defensive value and the misuse questions.

Why It Matters

Large 4 is a significant data point in the open-weight AI race. Over the past year, the strongest freely downloadable models have overwhelmingly come from Chinese labs such as DeepSeek, Zhipu's GLM and Moonshot's Kimi. A European trillion-parameter model trained on European soil gives governments and regulated companies in the EU a sovereign alternative they can host themselves, a selling point Mistral has leaned on with more than 125 enterprise clients including Airbus, ASML and HSBC.

For researchers, the release is a test of whether a well-funded lab outside the US-China duopoly can stay near the frontier. For practitioners, the practical calculus is straightforward: at roughly $1.36 per million input tokens, Large 4 undercuts many closed frontier models, and once weights ship it can run on private infrastructure. Developers Digest advises teams planning local or private-cloud deployments to wait for the weights and the final licence terms before committing hardware.

There is also a strategic message for the wider ecosystem. Mistral trained Large 4 from scratch rather than distilling from a larger proprietary system, and it did so on its own European compute. That demonstrates that a lab outside the largest US hyperscalers can still assemble the infrastructure needed for trillion-parameter training, an argument European policymakers have been eager to see validated.

What to Watch

  • October 27 weights drop: the licence terms will decide how freely companies can fine-tune and commercialise the model.
  • Independent benchmarks: third-party leaderboards will show whether the coding and cyber leads hold up.
  • Serving economics: a 1T-parameter MoE still needs substantial GPU memory, so expect quantised variants and hosting offers from cloud partners.

Sources