The Voice of Business

Mistral Large 4 review: Is it Right for your Business?

Mistral Large 4 is a 1-trillion-parameter model from France, with open weights promised for 27 October. What it claims, what is confirmed and who it suits.

A glossy black and purple pixel-block M logo glowing on a dark background beside the headline Mistral Large 4 review: Is it Right for your Business?

Mistral AI released Mistral Large 4 on 6 October 2026, two days before this review. It is the French company’s first new model since May, it is nicknamed “Le Chonk”, and it comes with a promise that matters to a lot of European businesses: the weights will be published for anyone to download. Those weights are not out yet, and most of the benchmark numbers come from Mistral itself. This review covers what the model is, what has been confirmed, what is still disputed, and whether it is right for your business today.

The short version: Mistral Large 4 is interesting mainly for where it comes from and how it will be released, and not yet for being the best model at anything most businesses do.

Key takeaways

  • It is a model of about 1 trillion parameters, with 52 billion active for each token. It accepts images and text and replies in text.
  • It is a public preview. The weights are promised for 27 October and the licence is not final.
  • Mistral’s best results are in cybersecurity: 93% on Cybench and 82% on a patching test. These are Mistral’s own numbers, though an independent check of a related test came out at 81.7%.
  • Coding is its weak side. It scored about 62% on DeepSWE v1.1 against about 74% for GPT-6 Astra on the live leaderboard.
  • Early coverage disagreed on basic facts, such as the number of active parameters and the context window. Mistral’s own pages settle some of them, but the licence is still unseen, so treat the figures as provisional.
  • It was trained in Europe and supports more than 160 languages, which is the strongest reason for a European business to look at it.

What Mistral Large 4 is

Large 4 is a large language model built as a sparse mixture of experts. That means the model holds about a trillion parameters but only uses 52 billion of them on any one token, which keeps the cost of running it far below what the headline size suggests. It is multimodal on the way in, since it reads images, but VentureBeat reports that Mistral’s research chief confirmed it still produces text only.

Mistral's model page for Mistral Large 4, marked public preview and open, describing 52B active and 1.05T total parameters, with a 1M token context
Source: Mistral's own model page for Large 4, on launch day.

Mistral built it from scratch over roughly two months on Nvidia Grace Blackwell chips in Mistral’s own European data centres. Mistral’s announcement says 3,800 of them, where Axios reported 4,000. Pierre Stock, Mistral’s vice president of science, told TechCrunch this is two to three times fewer than its Chinese competitors use. Mistral says a large share of the training data covers more than 160 languages, including every official EU language.

It follows Mistral Large 3, which had 675 billion parameters in total and 41 billion active. Mistral also expects larger and more capable versions in the coming months, and a later series of models tuned for specific jobs and built from Large 4. If you know Mistral from its chat assistant, see our entry on Le Chat.

Want this for your team?

Pick a time to talk it through with one of our trainers.

What it does well

Cybersecurity. This is where Mistral makes its strongest claim. It reports 93% on Cybench, a set of 40 security challenges, and 82% on a test that asks the model to reproduce and patch software vulnerabilities. Mistral says leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score close to zero on that second test, partly because they decline the task. Artificial Analysis, an independent aggregator, measured 81.7% on a closely related version of the test, which lends the number some weight. It is still a narrow result, and it says more about a willingness to do security work than about general ability.

Human judgement on coding. In a blind evaluation by Surge AI, annotators rated outputs on a scale of 1 to 5. Large 4 Preview scored 3.74, second of five models, behind Claude Opus 5 on 4.22 and ahead of Kimi K3 on 3.59 and GLM-5.3 on 3.60. That is Mistral’s reported result, but it is a better sign than the automated coding benchmarks suggest.

Legal and financial tasks. On the Harvey Legal Agent Benchmark Large 4 reached a 15% task pass rate, against 12.92% for Kimi K3 on the public Vals.ai board. That sounds modest, and VentureBeat points out that several closed models score higher on the same board. Vals.ai also ran a comparison commissioned by Mistral, which reportedly has Large 4 ahead of GPT-6 Astra on legal and financial tasks. A comparison the vendor paid for is a claim and not an independent finding.

Languages and Europe. Support for more than 160 languages, trained on European infrastructure, is a practical advantage if your customers or staff work in several European languages. SiliconANGLE also reports that Mistral describes computer vision, such as finding objects in images, as a strength.

Open weights, once they arrive. Pierre Stock argues that an open-weight model is easier to audit. For a regulated business that wants to inspect or host a model under its own rules, that is a real attraction, and it is the reason this launch has drawn more attention than a typical model update.

Where it falls short

Coding. Mistral says plainly that Large 4 trails frontier models here. On DeepSWE v1.1 it scored about 62%. The live leaderboard puts GPT-6 Astra, Claude Opus 5 and Gemini 3.8 Flash at around 74% and GLM-5.3 and Kimi K3 at around 69%. Eagentix also lists 28.3% on Terminal-Bench 4.0 and concludes it is a weaker fit for general-purpose coding agents. VentureBeat adds that the result “does not establish an outright coding lead”.

Overall ranking. On Artificial Analysis’s Intelligence Index, Large 4 Preview scores 38.4, which puts it eighth among open models. The seven above it are all Chinese. That does not contradict Mistral, because its claim is that it is the best open-weight model made outside China, and the previous best from outside China scored 33.6. It does mean “best open model” and “best open model outside China” are different claims.

The best claim is provisional. VentureBeat calls the “strongest open-weight model outside China” claim provisional until outside testers see the final model and weights. Mistral itself expects the benchmark results to change before release, and the preview is also being used for further tuning.

Hallucinations remain possible. A bigger model does not remove the need to check output. A confident wrong answer, what people call a hallucination, is still a risk with any model.

The claims that are still unconfirmed

Early coverage of this launch disagreed with itself, so it is worth knowing which numbers Mistral’s own pages settle and which they do not.

  • Active parameters. Mistral’s announcement and model page say 52 billion. Several outlets, and Mistral’s research page, said 49 billion.
  • Context window. Mistral’s model page lists 1 million tokens. Eagentix reported 512,000 from the announcement and about 520,000 measured by Artificial Analysis, so test long inputs before you rely on them.
  • Training hardware. Mistral says 3,800 chips. Axios reported 4,000.
  • Total size. The model page says 1.05 trillion parameters. Most outlets round it to 1 trillion.
  • Benchmarks at launch. TechCrunch said Mistral’s benchmark results were still pending, while Mistral and other outlets published figures the same day, and Mistral says the results may change before release.
  • Licence. The model page calls it open-weight but names no licence. VentureBeat says a custom Mistral licence is expected. The terms decide whether your business can use it commercially.

None of this makes the launch less real. It means the facts will keep moving until the weights are out and independent testers have had a go.

Price and access

The headline of Mistral's launch post, Le chonk: Introducing Mistral Large 4, on a dark background
Source: Mistral's launch post, dated 6 October 2026.

Large 4 is available as a public preview through Mistral’s API and its Mistral Studio platform, with a moderated version for everyone. Mistral is also red-teaming it with cybersecurity leaders, vetted partners and state authorities, who get a version with less moderation and broader cyber capability. Runtimewire notes the other side of open weights: once anyone can download and modify them, safeguards are easier to remove.

On price, the API is billed per million tokens, with separate rates for input, cached input and output. Mistral’s model page lists the rates and shows a 50% preview discount with no stated end date, so the price may change. This review does not quote the figures: check Mistral’s own pricing page. Per-token and per-task costs are also different measures. One comparison table in the Eagentix write-up shows Large 4 costing more per task than DeepSeek V4 Pro, even though it is cheaper than some of the other Chinese models.

What open weights mean for your Business

“Open weight” means you can download the finished model and run it yourself. It does not mean open source in the strict sense, because the training data and method may stay private and the licence may restrict use.

For a small or mid-sized business the practical effects are these. Mistral says the model can run on a private cloud or on premises, and that a European deployment will be operated end to end by Mistral under European law. Choosing where the model runs can simplify questions under GDPR. You could inspect and test it, and you are not locked to one vendor’s prices. Mistral’s enterprise pitch also mentions sovereign infrastructure and zero-data-retention options.

The catches are real. A trillion-parameter model needs serious hardware, so for most small businesses the sensible route is a hosted service, not their own machine. The weights are not published, the licence is unseen, and a European-trained model is not automatically compliant with anything. Where the data goes depends on the service you choose, so ask the provider.

How it compares

The chart below sets Large 4 against rivals on one coding test. The figures come from different sources and agent setups, so read them as a rough picture and not a league table.

Coding (DeepSWE v1.1). Large 4 trails the leaders.

On the other tests, the comparison is lopsided because Mistral chose them. Cybersecurity and some legal and financial work favour Large 4. General reasoning and coding favour the leading closed and Chinese models. If you are choosing between assistants for everyday office work, our reviews of Claude against ChatGPT and a local model against Claude are closer to your decision.

Who it suits

Large 4 suits businesses that care where their AI comes from and how it is run: regulated firms, public bodies, and anyone with a policy to prefer European suppliers. It also suits security teams, who may find it more willing to help with defensive work than models that decline, and teams that work in many European languages.

It suits others less. If your main use is writing code, you already have better options. If you want a polished assistant with apps and a mobile client, the big closed products are ahead. And if you were hoping to download and run it this week, you cannot yet.

Whatever you use it for, keep a person in the loop. A human in the loop who reads the result before it is final turns a risky tool into a useful one.

Verdict

Mistral Large 4 is not the best model in the world, and Mistral does not quite claim it is. It is a credible, European-built, soon-to-be open-weight model with a clear specialism in security and a pitch built around control. That is enough to make it worth watching, particularly on 27 October when the weights and licence should finally be public.

For most businesses the sensible move is to wait for those weights, the final licence and independent testing, then run one task you actually do on Large 4 and on the model you use now, and compare the results yourself. Treat every benchmark in this review, and especially the ones Mistral published, as a hint and not a verdict.

Further reading

Sources

FAQ

Questions we get asked

What is Mistral Large 4?

Mistral Large 4, nicknamed Le Chonk, is the flagship model from the French AI company Mistral AI, released as a public preview on 6 October 2026. It has about 1 trillion parameters in total, with 52 billion active for each token it processes. It accepts images as well as text but replies in text only.

Are the Mistral Large 4 weights available yet?

Not yet. Mistral says it will publish the weights by the end of October, and Reuters, as cited by other outlets, gives 27 October. Until then the model is only available through Mistral's own platform. The licence has not been finalised either, and VentureBeat reports it is expected to be a custom Mistral licence.

How much does Mistral Large 4 cost?

Through the API it is billed per million tokens, with separate rates for input, cached input and output. Mistral's model page lists the rates and shows a 50% preview discount with no stated end date, so the price you pay may change. Check Mistral's own pricing page before you budget.

Is Mistral Large 4 better than GPT-6 Astra or Claude?

Not across the board. Mistral's strongest results are in cybersecurity, and it reports that it beats GPT-6 Astra on a few narrow tests. On coding it trails the leading models, and on Artificial Analysis's overall index it ranks eighth among open models, behind seven Chinese ones.

Can I run Mistral Large 4 on my own servers?

Not today, because the weights are unpublished. Once they are out, a model with a trillion parameters will still need serious hardware, so most small businesses would use it through a hosted service rather than run it themselves.

Ready to Put This to Work?

Tell us where your team is with AI and we will tell you honestly what would make the biggest difference.

More comparisons

All comparisons