🔍 Read the full analysis: Mistral Large 4 For AI Agents: What Works And What Doesn’t on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 is available as an API research preview, with a 38.4 score on Artificial Analysis Intelligence Index v4.3.2. The source’s analysis credits a major improvement over Mistral’s earlier models but flags weaker benchmark results than leading US and several Chinese models, high output-token use, cost concerns and observed hallucinations. The model’s weights and licence were not yet available in the supplied material.
Mistral released Large 4 as a research preview on its API, and independent benchmark figures cited by ThorstenMeyerAI.com put it at 38.4 on Artificial Analysis Intelligence Index v4.3.2. The release marks a sharp improvement over the company’s previous models, but the analysis says Large 4 trails leading US models and several Chinese models on a benchmark weighted toward agentic tasks, raising questions about its cost and suitability for long-running AI agents.
Artificial Analysis’ index score for Large 4 is 38.4. In the same version of the index, the source lists US models including Claude Opus 5.5 at 57.6 and Gemini 4 Argon at 52.6, while Chinese models GLM-5.3 at 44.8 and Kimi K3 at 43.6 also rank above it. The source says Large 4 is the highest-scoring model from outside the United States and China, while cautioning that this comparison reflects a limited competitive field.
The model has one trillion parameters, with 49 billion active, accepts text and images, produces text, and has a 512,000-token context window. It is currently a proprietary API preview. The source says Mistral promised to release weights by the end of October, but reports that the licence had not been published. Mistral also said reinforcement learning was ongoing, so scores could change.
At listed rates, the API costs $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. The source reports a two-week introductory discount of 50%. It also cites Artificial Analysis data showing Large 4 used 200 million output tokens to complete the index, compared with a median of 81 million for comparable models. That difference can affect both spending and response time in agent workflows.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
Agent Workloads Face Cost and Reliability Tests
The index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. That makes its score relevant to buyers evaluating models for multi-step knowledge work, software tasks and SaaS workflows, though a benchmark result cannot establish how a model will perform in every deployment.
The source’s analysis argues that capability gaps may matter more on extended tasks than on short exchanges: an error early in a sequence can shape later actions. It also flags output volume as a practical concern. If an agent produces substantially more tokens while completing comparable work, cost and latency can accumulate across steps, even when the advertised per-token rate appears manageable.
The article also reports that its author observed confident false statements in hands-on use. That is an attributed personal observation, not a published benchmark finding. For an agent, an unsupported claim can become an input to later decisions, so buyers should test factual reliability and require suitable checks before assigning consequential tasks.
As an affiliate, we earn on qualifying purchases.
A Large Jump From Mistral’s Earlier Scores
The same Artificial Analysis index version gave Mistral Large 3 a score of 9 and Medium 3.5 a score of 14, according to the source. Large 4’s 38.4 therefore represents a substantial advance on those earlier results. The source describes it as an unusually large step for a European lab, while noting that the gain does not put it among the top-scoring US or Chinese models.
The source’s comparisons also show why the launch framing needs qualification. Large 4 scores above some earlier Chinese models, including GLM-5.2 at 33.7 and DeepSeek V4 Pro at 36.0, but below newer entries such as GLM-5.3 and DeepSeek V4.1 Flash. Artificial Analysis figures are benchmark results, not a universal measure of model quality; performance depends on the task, setup and evaluation method.
“Reinforcement learning is still running.”
— Mistral, as reported by ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
Weights, Licensing and Reliability Remain Open
The supplied source says the model weights were promised for the end of October, but does not provide a calendar date for the release announcement or confirm that the weights were subsequently published. It also says the licence was unpublished at the time of writing. Until those details are available, developers cannot assess the terms for self-hosting or reuse from the source material alone.
Benchmark scores may shift because Mistral said reinforcement learning was continuing. The source does not provide a full account of its hands-on test method, sample size or the specific prompts behind the reported hallucinations. Nor does the supplied material establish how Large 4 performs across independent production deployments. Buyers should treat the author’s observation and the benchmark data as different kinds of evidence.
As an affiliate, we earn on qualifying purchases.
Watch for Weights and Updated Evaluations
The next milestones are publication of the model weights and licence, along with any updated Artificial Analysis results after further reinforcement learning. Those developments would clarify whether the model can be run outside Mistral’s API and whether its benchmark standing changes.
For now, teams considering Large 4 can compare it on their own agent tasks against alternatives, tracking completion rates, factual errors, token consumption, latency and total cost. The supplied source supports neither a blanket recommendation nor a conclusion that the model is unsuitable for every use. Its evidence instead points to a clear need for task-specific testing before production adoption.
AI agent performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Mistral Large 4?
Mistral Large 4 is a multimodal model available as a research preview through Mistral’s API. The source describes it as accepting text and images, producing text, and supporting a 512,000-token context window.
How did it score against other models?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, as cited by the source. Several listed US and Chinese models scored higher; results are specific to this index version and should not be treated as a complete measure of performance on every task.
Is Large 4 open source?
Not according to the supplied source at the time of writing. It describes Large 4 as a proprietary API preview, with weights promised for the end of October and the licence not yet published.
What should agent developers test before using it?
Teams should measure task completion, factual reliability, output-token use, latency and total cost on their own workflows. The source reports high benchmark token use and an author-observed hallucination issue, but does not establish performance across all deployments.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
