🔍 Read the full analysis: Where Mistral Large 4 Shines—and Why It May Not Suit Your Agents on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on Artificial Analysis’ Intelligence Index, a sharp improvement over the company’s earlier models but below leading US and Chinese systems. The supplied review flags its benchmark-task cost, high output volume and reported confident errors as concerns for multi-step agents; the model remains in research preview, with weights and licence details pending.
Mistral released Large 4 as a research preview, and results cited from the Artificial Analysis Intelligence Index v4.3.2 show a substantial jump from the company’s previous models. Its score of 38.4 remains below leading US and Chinese systems, while the source report raises concerns about cost, output length and observed hallucinations for buyers considering it for multi-step agent tasks.
Large 4 is a 1 trillion-parameter model with 49 billion active parameters, according to the supplied report. It accepts text and images, produces text, and has a 512,000-token context window. Mistral offers it through its API as a research preview; the company has promised to release model weights at the end of October. Until then, the weights are unavailable and the licence has not been published, the report says.
On Artificial Analysis’ current index, Large 4 scored 38.4. The report compares that with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version, describing the result as a steep improvement. The score is nevertheless below the listed US leaders, headed by Claude Opus 5.5 at 57.6, and below several Chinese models, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5. The report says the ranking places Large 4 eighth among open-weight models if and when its weights are released.
The report lists API pricing at $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; a 50% discount applies for the first two weeks. It calculates Large 4’s cost at $1.13 per Intelligence Index task. By comparison, it gives $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, which both score higher on the index. Artificial Analysis figures and the report’s calculations reflect the stated benchmark and pricing assumptions, not every workload.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The Trade-Off for Agent Work
The result matters to teams choosing a model to carry out long-running, multi-step work, not just answer short prompts. The index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. A lower overall score does not translate directly into a particular failure rate, but weaknesses can compound as an agent takes actions and builds later steps on earlier outputs.
Cost is also tied to how much a model writes, not only its listed price per token. The source report says Large 4 used 200 million output tokens across the index, compared with a median of 81 million for comparable models. That is a benchmark observation; an individual customer’s token use will depend on its prompts, tools and tasks. Still, higher output volume can add expense and latency in workflows that call the model repeatedly.
The report’s author says hands-on testing found instances of confidently stated errors. That is a personal observation, not a result from the Artificial Analysis index, and no test method or error rate is provided in the supplied material. It is a relevant warning for agents, where an incorrect claim may shape subsequent actions, but it is not enough on its own to establish how often that will happen in production.
As an affiliate, we earn on qualifying purchases.
From Mistral 3 to Large 4
The index comparison gives a measure of Mistral’s recent progress: the supplied report lists Large 3 at 9 and Medium 3.5 at 14, against Large 4’s 38.4 on the same version. That makes the new score a significant step for the company, even though it does not put the model among the highest-scoring systems listed.
The headline framing in the source material describes Large 4 as the most intelligent model from outside the United States and China. The listed index results support a narrower point: it is a strong result for a French model, while the cited US and Chinese models score higher. Such regional comparisons depend on which models are included and how they are evaluated; they do not by themselves establish that Large 4 is the best choice for a particular task.
Mistral says reinforcement learning is still underway and that scores may change, according to the supplied report. That makes the current benchmark a preview-stage snapshot, rather than a final assessment of the released weights or their performance in customer deployments.
“Large 4 shows it has closed a lot of ground — and that it is still not one.”
— ThorstenMeyerAI.com report
As an affiliate, we earn on qualifying purchases.
Preview Results and Open Questions
Several details remain unsettled. The supplied material does not give a calendar date for the release, only that it happened the day before the report, and Mistral’s promised end-of-October weights release has no year specified. The weights’ final contents and licence terms are not yet available in the source.
The index score may move as Mistral continues reinforcement learning. The report also does not provide a reproducible method, sample size or error rate for its hands-on hallucination observations, so readers cannot use them as a quantified comparison with other models. Nor do benchmark task costs predict the bill for every real-world agent: workload, prompt design, caching and output length all affect usage.
AI model performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Weights, Licensing and Retesting
The next stated milestone is Mistral’s planned release of Large 4 weights at the end of October. The company has not published a licence in the supplied material, so prospective users will need to review the terms when they become available before treating the model as an open-weight option.
Benchmark scores may also change as reinforcement learning continues. Buyers evaluating Large 4 for agents can compare updated independent results with tests on their own tasks, tracking completion rates, errors, token use and latency. Until those results and deployment terms are clearer, the current evidence supports a notable Mistral performance gain, but not a blanket recommendation for agent workloads.
AI output hallucination detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Mistral Large 4?
It is Mistral’s 1 trillion-parameter text-and-image input model, with 49 billion active parameters and a 512,000-token context window, offered through an API research preview, according to the supplied report.
How did Large 4 score against leading models?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The listed US leaders and several Chinese models scored higher; the report says the result is a major improvement on Mistral’s earlier index scores.
Why does the report question its use for agents?
The report points to its lower benchmark score than leading systems, higher output-token use and the author’s own unquantified observations of confident errors. These concerns may matter more in workflows where mistakes can carry through several steps.
How much does Large 4 cost?
The report lists API rates of $1.36 per million input tokens and $4.18 per million output tokens, plus $0.14 per million cached input tokens. It says a 50% discount applies for the first two weeks; actual costs vary with usage.
When will the model weights be available?
Mistral has promised weights for the end of October, according to the source material. It does not specify the year there, and the licence has not yet been published in the supplied information.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
