An open-weight pitch with no weights yet

Mistral is selling Mistral Large 4 as an open-weight model, but the only license terms it has published for it are proprietary. On 6 October 2026, Mistral launched a public preview of the model, nicknamed Le Chonk, on its paid API. The post says ML4 "pushes the frontier of open-weight performance" and that "Weights drop end of this month."

Mistral's model page lists 1.05T total parameters, 49B active and a 1.6B vision encoder. It tags the model Open and Public Preview, while its weights tab shows the license and the weights as coming soon.

When The Frontier checked the mistralai account on Hugging Face at 23:18 SGT on 6 October, it held no Large 4 repository. Its newest model was Shieldstral 1.0 3B, created on 16 July 2026. Mistral Large 3 sits there under Apache 2.0.

Mistral cofounder Guillaume Lample wrote on X that the RL run behind the preview "is still in flight". He added: "we will release a final version before the end of the month along with the weights of the model."

Reflection made a similar promise for its Beam model on 5 October. See Reflection previews Beam, a 501B open-weight MoE model with 23B active, and says Apache 2.0 weights come this month.

Mistral's AI Act papers list proprietary terms

Mistral's legal center entry for Mistral Large 4 classifies it as a general purpose AI model, released on 6 October 2026.

The technical documentation for downstream providers says it will be updated in case of an open weights release. Until then it names three sets of terms:

  • The Terms of Service and the partner terms are each described as a proprietary license that restricts access to the weights.
  • A confidential self-deployment agreement is described as a bespoke license that allows access to the weights, with possible limits on use and modification.

The same document's size table gives 250 tokens as the maximum text input and output. The docs page lists 1M tokens of context, while Artificial Analysis and Vals AI both list 512k.

The public summary of training content ticks the box for more than 10 trillion text tokens. Common Crawl is the only dataset it names, and it says data from users of other Mistral products was used. It confirms that crawlers were used, but answers NA where the template asks for their names.

What independent tables show beside Mistral's charts

Independent scores arrived on launch day. On the Artificial Analysis Intelligence Index, the preview scores 38. AA scores MiMo-V2.6-Pro at 46, GLM-5.3 at 45 and Kimi K3 at 44. GLM-5.3-Flash scores 42 and DeepSeek V4.1 Flash 39, and AA labels all five open weights. AA's full model list also scores two open-weight Alibaba models, Qwen3.8 2.4T A95B and Qwen3.8-Flash-Next, at 40. AA labels Mistral Large 4 itself a proprietary model for now.

Mistral's X post limits its ranking claim to two regions, calling ML4 "the best open weights model from US or Europe on aggregated benchmarks." Artificial Analysis wrote on X that France "is back to having the most intelligent model from outside the US and China".

Mistral bar chart of Vals AI Finance Agent v2 scores for Mistral Large 4 Preview and seven other models Mistral's own Finance Agent v2 chart shows GLM-5.3 at 55.8 and Mistral Large 4 Preview at 54.7.
Source: Mistral, Introducing Mistral Large 4, 6 October 2026

Mistral says ML4 is state of the art among open models on finance. Its own chart above puts GLM-5.3 ahead, and the Vals AI model page ranks ML4 22nd of 75 on Finance Agent v2. Vals ranks it sixth of 75 on Harvey's Legal Agent Benchmark, at 15.83%.

On cyber, Mistral says ML4 ranks among the top five models on the Artificial Analysis Cyber Index. AA's page puts Grok 4.7 and MiMo-V2.6-Pro at 56 and GPT-6 Luna at 53, ahead of ML4 and GLM-5.3-Flash at 50. Neither of the two Cyber Index charts in Mistral's post shows Grok 4.7, MiMo-V2.6-Pro or GPT-6 Luna. Anthropic assessed the cyber skills of open-weight GLM-5.3 on 30 September. See Anthropic says Z.ai GLM-5.3 brings Mythos-class cyber skills with weak open-weight safeguards.

Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4, plus a Coding Agent Index of 49.8%. The 49.8% figure equals the simple average of those three scores. Mistral says the index puts ML4 ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. Its own charts show Kimi K3 at 68 on DeepSWE and GLM-5.3 at 40 on Terminal-Bench 4.

Those two charts say AA evaluated the scores privately before the harness launched. AA's public model page now shows about 27% for ML4 on Terminal-Bench 4.0, and Vals shows 22.73%.

Vals puts ML4 32nd of 44 on its Vals Index at 48.05%, costing $13.78 per test. The Vals page for GLM 5.3 shows 53.51% at $7.25 per test. A Hacker News commenter posting as tosh wrote: "sorting the charts like that gives off weird vibes".

Where independent leaderboards place Mistral Large 4

Artificial Analysis and Vals AI were the only public leaderboards The Frontier found listing ML4 between 23:24 and 23:30 SGT on 6 October. LMArena, SWE-bench, Scale's SEAL and SWE-bench Pro boards, Terminal-Bench, Epoch AI, LiveBench, Aider, Stanford's HELM and OpenRouter's rankings did not.

Mistral says ML4 leads open-weight models developed outside China on the AA Cyber Index "by a wide margin." NVIDIA's Nemotron 3 Ultra, the only US open-weight model in that index, scores 13 against ML4's 50. On AA's CyberGym-E2E-AA table, ML4 leads every model at 82%, with MiMo-V2.6-Pro at 79%.

Mistral says ML4 is state of the art among open-weight models on SciCode-Verified. On AA's SciCode table, open-weight MiMo-V2.6-Pro scores 61%, Kimi K3 60% and GLM-5.3 59%, against 54% for ML4.

Mistral also says ML4's 59.9% on AutomationBench puts it ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro. AA's AutomationBench-AA table agrees on those three. It also scores open-weight DeepSeek V4.1 Flash at 68.9%, MiMo-V2.6-Flash at 64.1% and GLM-5.3 at 62.2%.

On AA-Briefcase, AA rates ML4 at 1,393, above DeepSeek V4 Pro 0813 at 1,256. MiMo-V2.6-Pro rates 1,516, GLM-5.3 1,510 and Kimi K3 1,501. AA's Terminal-Bench 4.0 table shows ML4 at 27% and Kimi K3 at 13%, as Mistral's Baptiste Rozière wrote on X, with GLM-5.3 at 42%.

Vals AI labels ML4 open weight, unlike AA. On the Vals Index, open-weight MiMo V2.6 Pro scores 55.20%, GLM 5.3 53.51% and DeepSeek V4.1 Flash 51.32%, against 48.05% for ML4. Thinking Machines' open-weight Inkling scores 28.67%.

Mistral's post says ML4 significantly outperforms any open-weight model developed in the US or Europe. Inkling scores higher on two Vals tables, MedCode at 41.19% against 40.69% and MedScribe at 85.41% against 80.43%. On AA's long context reasoning test, Meta's open-weight Muse Glimmer scores 83% to ML4's 81%.

On finance, Vals scores open-weight GLM 5.3 Flash at 57.85% and MiMo V2.6 Pro at 57.34% on Finance Agent v2. On Harvey's Legal Agent Benchmark, Vals says ML4 trails five models, all of which it marks proprietary, led by Muse Spark 1.2 at 25.42%. On Vals's Legal Research Bench, ML4 scores 31.73%, below GPT-6 Astra at 39.42% and open-weight GLM 5.3 at 49.04%.

The AI Act exemption needs published weights

Article 53 of the AI Act exempts some open models from its documentation duties toward the AI Office and downstream providers. The European Commission's questions and answers on general-purpose AI models tie the exemption to a free and open-source license. The weights, architecture and usage information must also be public.

That exemption does not apply to models with systemic risk. The Commission's page sets a training compute threshold of 10^25 FLOP and says providers must notify it within two weeks of meeting it.

Mistral is listed as a signatory of the General-Purpose AI Code of Practice, and its training summary says so too. Neither Mistral document for ML4 gives a training compute figure.

Vendor charts and safety scores we set aside

Several of Mistral's charts carry Artificial Analysis or Vals AI scores, but Mistral chose which rivals each chart shows. Its blind coding evaluation with Surge AI ranked ML4 second of five at 3.74, behind Claude Opus 5 at 4.22. Anthropic's newer Opus 5.5 was not in that test.

Mistral's safety results are self-reported. It cites 93.3% attack resistance on Lakera's B3 benchmark, where its own chart shows GLM-5.3 tied at 93.3. Its B3, KORA and cyber refusal charts compare ML4 only with GLM, Kimi and DeepSeek models. It says red teamers will get "reduced moderation and expanded cyber capabilities" before the weights ship.

AA's thread reports the same 82% on its CyberGym-E2E-AA test that Mistral cites. Mistral says Claude Opus 5.5 and GPT-6 Astra score near zero there because they refuse the task.

Weights, license and compute still missing

This article could not find:

  • weights, a license text or a model card on Hugging Face
  • a system card or any safety evaluation run by a third party
  • a training compute figure beyond the 3,800 Grace Blackwell GPUs the post names
  • the crawler names the AI Act template asks for

Mistral lists Large 4 at $1.36 input and $4.18 output per million tokens, and the docs page now shows $0.68 and $2.09 beside those rates. AA says that 50% discount lasts two weeks, and no Mistral page we read gives an end date. Mistral Large 3 costs $0.50 and $1.50, and Medium 3.5 costs $1.50 and $7.50. Large 4's list rates are 2.7 times Large 3's for input and 2.8 times for output, and 44% below Medium 3.5's for output.

Mistral's model lifecycle policy says Public Preview models allow silent updates and may be retired before general availability. Scores published this week therefore describe a preview checkpoint that Mistral's own policy lets it change without notice.