AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Where Mistral Large 4 Shines—and Why It May Not Suit Your Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis’ Intelligence Index, a sharp improvement over the company’s earlier models but below leading US and Chinese systems. The supplied review flags its benchmark-task cost, high output volume and reported confident errors as concerns for multi-step agents; the model remains in research preview, with weights and licence details pending.

Mistral released Large 4 as a research preview, and results cited from the Artificial Analysis Intelligence Index v4.3.2 show a substantial jump from the company’s previous models. Its score of 38.4 remains below leading US and Chinese systems, while the source report raises concerns about cost, output length and observed hallucinations for buyers considering it for multi-step agent tasks.

Large 4 is a 1 trillion-parameter model with 49 billion active parameters, according to the supplied report. It accepts text and images, produces text, and has a 512,000-token context window. Mistral offers it through its API as a research preview; the company has promised to release model weights at the end of October. Until then, the weights are unavailable and the licence has not been published, the report says.

On Artificial Analysis’ current index, Large 4 scored 38.4. The report compares that with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version, describing the result as a steep improvement. The score is nevertheless below the listed US leaders, headed by Claude Opus 5.5 at 57.6, and below several Chinese models, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5. The report says the ranking places Large 4 eighth among open-weight models if and when its weights are released.

The report lists API pricing at $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; a 50% discount applies for the first two weeks. It calculates Large 4’s cost at $1.13 per Intelligence Index task. By comparison, it gives $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, which both score higher on the index. Artificial Analysis figures and the report’s calculations reflect the stated benchmark and pricing assumptions, not every workload.

At a glance
analysisWhen: Released the day before the source repo…
The developmentMistral released Large 4 as a research preview, prompting an assessment of its benchmark gains, costs and suitability for agent workflows.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Trade-Off for Agent Work

The result matters to teams choosing a model to carry out long-running, multi-step work, not just answer short prompts. The index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. A lower overall score does not translate directly into a particular failure rate, but weaknesses can compound as an agent takes actions and builds later steps on earlier outputs.

Cost is also tied to how much a model writes, not only its listed price per token. The source report says Large 4 used 200 million output tokens across the index, compared with a median of 81 million for comparable models. That is a benchmark observation; an individual customer’s token use will depend on its prompts, tools and tasks. Still, higher output volume can add expense and latency in workflows that call the model repeatedly.

The report’s author says hands-on testing found instances of confidently stated errors. That is a personal observation, not a result from the Artificial Analysis index, and no test method or error rate is provided in the supplied material. It is a relevant warning for agents, where an incorrect claim may shape subsequent actions, but it is not enough on its own to establish how often that will happen in production.

Amazon

AI model API key management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Mistral 3 to Large 4

The index comparison gives a measure of Mistral’s recent progress: the supplied report lists Large 3 at 9 and Medium 3.5 at 14, against Large 4’s 38.4 on the same version. That makes the new score a significant step for the company, even though it does not put the model among the highest-scoring systems listed.

The headline framing in the source material describes Large 4 as the most intelligent model from outside the United States and China. The listed index results support a narrower point: it is a strong result for a French model, while the cited US and Chinese models score higher. Such regional comparisons depend on which models are included and how they are evaluated; they do not by themselves establish that Large 4 is the best choice for a particular task.

Mistral says reinforcement learning is still underway and that scores may change, according to the supplied report. That makes the current benchmark a preview-stage snapshot, rather than a final assessment of the released weights or their performance in customer deployments.

“Large 4 shows it has closed a lot of ground — and that it is still not one.”

— ThorstenMeyerAI.com report

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results and Open Questions

Several details remain unsettled. The supplied material does not give a calendar date for the release, only that it happened the day before the report, and Mistral’s promised end-of-October weights release has no year specified. The weights’ final contents and licence terms are not yet available in the source.

The index score may move as Mistral continues reinforcement learning. The report also does not provide a reproducible method, sample size or error rate for its hands-on hallucination observations, so readers cannot use them as a quantified comparison with other models. Nor do benchmark task costs predict the bill for every real-world agent: workload, prompt design, caching and output length all affect usage.

Amazon

AI model performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Retesting

The next stated milestone is Mistral’s planned release of Large 4 weights at the end of October. The company has not published a licence in the supplied material, so prospective users will need to review the terms when they become available before treating the model as an open-weight option.

Benchmark scores may also change as reinforcement learning continues. Buyers evaluating Large 4 for agents can compare updated independent results with tests on their own tasks, tracking completion rates, errors, token use and latency. Until those results and deployment terms are clearer, the current evidence supports a notable Mistral performance gain, but not a blanket recommendation for agent workloads.

Amazon

AI output hallucination detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is Mistral’s 1 trillion-parameter text-and-image input model, with 49 billion active parameters and a 512,000-token context window, offered through an API research preview, according to the supplied report.

How did Large 4 score against leading models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The listed US leaders and several Chinese models scored higher; the report says the result is a major improvement on Mistral’s earlier index scores.

Why does the report question its use for agents?

The report points to its lower benchmark score than leading systems, higher output-token use and the author’s own unquantified observations of confident errors. These concerns may matter more in workflows where mistakes can carry through several steps.

How much does Large 4 cost?

The report lists API rates of $1.36 per million input tokens and $4.18 per million output tokens, plus $0.14 per million cached input tokens. It says a 50% discount applies for the first two weeks; actual costs vary with usage.

When will the model weights be available?

Mistral has promised weights for the end of October, according to the source material. It does not specify the year there, and the licence has not yet been published in the supplied information.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The unbundling of the budget app. Why a conversational finance surface absorbs what the personal-finance apps charge for, and what survives the absorption.

OpenAI launched a personal-finance feature inside ChatGPT, absorbing commodity layers of budgeting apps. The category splits, with high-friction services surviving separately.

The Blueprint For AI In A Canada-EU Union

A detailed analysis of the new Canada-EU AI collaboration, highlighting model capabilities, licensing differences, and strategic implications for the AI landscape.

Two Channels: How the Pentagon Just Split Frontier-AI Procurement in Half

The Pentagon has split its AI procurement into two separate channels, placing Anthropic exclusively in a cybersecurity-focused stream, not excluded from federal contracts.

OpenAI’s 722 AI Mathematics Proofs Prompt A Question About What’s Next

OpenAI says an unreleased model produced 722 mathematical manuscripts. Outside mathematicians have not yet confirmed the claimed results.