Mistral Large 4 is the first European model in a year that belongs in the same conversation as GLM 5.3 and DeepSeek V4 Pro, and it is priced almost exactly like them. Mistral AI launched it on October 6, 2026 as a public API preview: 1 trillion parameters, 49 billion active per token, $1.36 per million input tokens and $4.18 per million output tokens, with open weights promised by the end of the month.

I went through Mistral’s launch post and the early independent numbers. The short version for developers is that Large 4 is a strong repository-level coding model and a very strong cybersecurity model. On terminal-heavy agent work, the one number Mistral published is weak.

What Is Mistral Large 4?

Mistral Large 4 is a natively multimodal, hybrid instruct-and-reasoning mixture-of-experts model with 1T total and 49B active parameters, trained from scratch by Mistral AI. Mistral’s nickname for it is “Le Chonk”. CEO Arthur Mensch introduced it at AI Everything Abu Dhabi on October 6.

The facts from the launch post and early coverage:

  • Parameters: 1 trillion total, 49 billion active per token
  • Modalities: text and image input, text output
  • Training: 3,800 Nvidia Grace Blackwell GPUs in Mistral’s own European datacentres; The Next Web reports about two months of training at roughly 10 megawatts
  • Languages: 160+, including every official EU language
  • Price (preview): $1.36 input / $4.18 output per million tokens
  • Context: Mistral’s post does not state one; Artificial Analysis lists roughly 520K tokens
  • Weights: “by the end of the month”, with October 27 reported by VentureBeat and The Next Web

That missing context figure is a small thing, but it bothered me. A launch post with nine benchmark rows and no context window means you should check the API’s actual limits before you design around them.

Mistral Large 4 is a frontier-scale open-weight contender from a European lab, available today only as a hosted preview.

Mistral Large 4 Benchmarks for Coding

Mistral published agentic coding numbers that put Large 4 at the top of the open-weight pack on one benchmark and well behind it on another.

DeepSWE v1.1 score (vendor-reported)
  • Mistral Large 4 62%
  • GLM 5.3 61%
  • DeepSeek V4 Pro 57%
  • Qwen 3.8 Max 51%
  • Reflection Beam 44%

Source: Mistral AI preliminary results, via The Next Web (Oct 2026)

Mistral’s own post gives the precise figure as 61.7% on DeepSWE v1.1. That is a one-point lead over GLM 5.3, which is inside the noise for a benchmark like this. Call it a tie with the best widely available Chinese open model, and a clear lead over DeepSeek V4 Pro and Qwen 3.8 Max.

The newer cheap models change the picture, though. Reflection AI’s launch table from the day before, which I covered in the Reflection Beam benchmark breakdown, lists DeepSeek V4.1 Flash at 74.2 on DeepSWE v1.1. Mistral did not include V4.1 Flash in its comparison.

Here is the number most of the launch coverage skipped:

BenchmarkMistral Large 4What it measures
DeepSWE v1.161.7%Agentic software engineering on real repos
SWE-Atlas QnA59.4%Answering questions about a large codebase
AutomationBench59.9%Multi-step workflow automation
Terminal-Bench 4.028.3%Long agent sessions in a real terminal
Human eval, coding (1–5)3.74Surge AI annotators, blind; Claude Opus 5 scored 4.22

(Source: Mistral AI, “Introducing Mistral Large 4”, October 6, 2026. All scores vendor-reported.)

28.3% on Terminal-Bench 4.0 is a long way behind the closed models. Anthropic reports 66.4% for Claude Opus 5.5 and Claude Sonnet 5.5 is reported at 70.6% on the same benchmark. If your agent lives in a shell, running builds, reading logs and retrying, that gap is the one that will show up in your day.

The SWE-Atlas QnA score is the more encouraging row. Answering questions about a repository you have been inside for an hour is most of what a coding agent does, and 59.4% is respectable. Reflection Beam managed 34.6 on the related SWE Atlas Codebase QnA test.

On Mistral’s own numbers, Large 4 ties the best open models on repo-level coding but trails closed frontier models by more than 35 points on terminal agent work.

The Cybersecurity Numbers Are the Real Headline

Mistral spent more of its launch post on security than on coding, and the claims are striking.

Large 4 scores 93% on Cybench, which Mistral calls one of the highest results reported for an open-weight model. On the AA Cyber Index’s vulnerability reproduction and patching test it scores 82%. Mistral’s post then adds a line worth reading carefully: “Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task.”

That is not a capability comparison. It is a policy comparison. Opus 5.5 scoring near zero because it refuses tells you about Anthropic’s safety settings, not about what the model could do.

It is still a real difference for defenders. A security team that wants a model to reproduce a CVE and write the patch will get an answer from Large 4 and a refusal from the closed models. Co-founder Guillaume Lample framed it as “cyber defense capabilities” for enterprises and governments, according to The Next Web.

The open question is what happens when anyone can download the weights. Mistral says Large 4 resists 93.3% of attacks on the B3 AI security benchmark and has a higher cyber refusal rate than comparable open models. TechCrunch quotes Mistral VP of Science Pierre Stock saying the company will “work with trusted partners and governments to make sure that the open source weights can be used to defend, but not to [perform] malicious attacks.” How that works once the weights are public is not explained.

Large 4 is the most capable openly available model for offensive-security-adjacent defence work, and that is both its selling point and its biggest policy risk.

Mistral Large 4 Pricing vs GLM 5.3, DeepSeek and Claude

At $1.36 / $4.18, Mistral priced Large 4 in the same band as the Chinese open models, not as a premium European product.

ModelInput / 1MOutput / 1MDeepSWE v1.1
Mistral Large 4 (preview)$1.36$4.1861.7%
DeepSeek V4 Pro (Together AI)$1.32$3.9657%
GLM 5.3 (Together AI)$1.40$4.4061%
Kimi K3 (Together AI)$2.70$13.5068.0
Claude Sonnet 5.5$2.00$10.00not reported

(Sources: Mistral AI launch post; Together AI pricing, October 2026; DeepSWE for GLM 5.3 and DeepSeek V4 Pro from Mistral’s comparison, Kimi K3 from Reflection AI’s launch table. All benchmark scores vendor-reported.)

DeepSeek V4 Pro is 3% cheaper on input and 5% cheaper on output, and Mistral’s own chart puts it 5 points lower on DeepSWE. GLM 5.3 is within pennies on both. Kimi K3 scores higher on DeepSWE in Reflection’s table, but its output price is more than three times higher.

Artificial Analysis measured 116.1 output tokens per second on Mistral’s API and an 18.69-second time to first token on its 10,000-token test workload. That first-token time includes reasoning, but it is long. For an interactive coding assistant, an 18-second pause before anything appears is noticeable. For a background agent, it does not matter much.

If you track these numbers across vendors, the AI model pricing tracker has the current per-token prices with sources.

Mistral Large 4 does not undercut the Chinese open models on price; it matches them, so the reason to pick it has to be something other than cost.

Who Should Use Mistral Large 4?

There are three groups where Large 4 is the obvious choice.

European teams with data residency rules. Mistral says its European deployment runs independently of its other regions. For a bank or public-sector team that cannot send code to a US or Chinese provider, Large 4 is now a credible frontier-scale option rather than a compromise.

Security teams. If your work involves reproducing vulnerabilities and writing patches, the closed models’ refusals are a real blocker, and Large 4 is built to do that work. Do your own evaluation on your own targets first. The 93% Cybench figure is vendor-reported.

Teams that need a non-Chinese open model for procurement. This is the same group I described in the Reflection Beam post. Large 4 scores 17 points higher than Beam on DeepSWE in Mistral’s chart, which makes it the stronger Western open option on that test once the weights land.

For everyone else, the decision is harder. If you are choosing an open model for coding this week, GLM 5.3 scores within a point on DeepSWE at about the same price, and its weights are already available. If you mainly need a cheap, fast model, DeepSeek V4.1 Flash costs $0.30 / $1.20 on Together AI.

You can also try it inside the agent you already use. The guide to running Claude Code with open models through Together Link covers how routing to non-Anthropic models works in practice, though Together AI does not list Large 4 yet.

Pick Large 4 for EU data residency, security research or Western-origin procurement rules; for general coding on a budget, GLM 5.3 is the equal-cost alternative you can self-host today.

What Is Still Missing Before You Build on It

Four things are unresolved as of October 7, 2026.

The weights are not out. Mistral says the end of October, and two outlets report October 27. Until then this is a hosted API only.

The license is unclear. Mistral’s announcement does not name a license. VentureBeat reports a custom Mistral license, and Artificial Analysis currently lists the preview as proprietary. “Open weights” under a custom license can still restrict commercial use, so read the terms when they arrive.

The benchmarks are preliminary. VentureBeat notes the results had not been independently verified on public leaderboards. Artificial Analysis rates the preview at 38 on its Intelligence Index, against a median of 26 for reasoning models in its price tier. That is a long way under Claude Opus 5.5 at 58, and it is a better reality check than any single vendor chart.

The preview will change. Mistral describes the next three weeks as a period of testing and further reinforcement learning. The model you call today may not be the one whose weights ship.

If you want the wider context on how open models compare with Claude and GPT for coding right now, the Claude Opus 5.5 vs GPT-6 comparison covers the closed side of the market.

Treat Mistral Large 4 as a model to evaluate in October and decide on in November, once the weights, license and independent benchmarks are public.