AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 For AI Agents: What Works And What Doesn’t on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 is available as an API research preview, with a 38.4 score on Artificial Analysis Intelligence Index v4.3.2. The source’s analysis credits a major improvement over Mistral’s earlier models but flags weaker benchmark results than leading US and several Chinese models, high output-token use, cost concerns and observed hallucinations. The model’s weights and licence were not yet available in the supplied material.

Mistral released Large 4 as a research preview on its API, and independent benchmark figures cited by ThorstenMeyerAI.com put it at 38.4 on Artificial Analysis Intelligence Index v4.3.2. The release marks a sharp improvement over the company’s previous models, but the analysis says Large 4 trails leading US models and several Chinese models on a benchmark weighted toward agentic tasks, raising questions about its cost and suitability for long-running AI agents.

Artificial Analysis’ index score for Large 4 is 38.4. In the same version of the index, the source lists US models including Claude Opus 5.5 at 57.6 and Gemini 4 Argon at 52.6, while Chinese models GLM-5.3 at 44.8 and Kimi K3 at 43.6 also rank above it. The source says Large 4 is the highest-scoring model from outside the United States and China, while cautioning that this comparison reflects a limited competitive field.

The model has one trillion parameters, with 49 billion active, accepts text and images, produces text, and has a 512,000-token context window. It is currently a proprietary API preview. The source says Mistral promised to release weights by the end of October, but reports that the licence had not been published. Mistral also said reinforcement learning was ongoing, so scores could change.

At listed rates, the API costs $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. The source reports a two-week introductory discount of 50%. It also cites Artificial Analysis data showing Large 4 used 200 million output tokens to complete the index, compared with a median of 81 million for comparable models. That difference can affect both spending and response time in agent workflows.

At a glance
analysisWhen: Reported as released the day before the…
The developmentMistral released Large 4 as an API research preview, prompting scrutiny of its agentic benchmark performance, costs and readiness for production workflows.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Workloads Face Cost and Reliability Tests

The index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. That makes its score relevant to buyers evaluating models for multi-step knowledge work, software tasks and SaaS workflows, though a benchmark result cannot establish how a model will perform in every deployment.

The source’s analysis argues that capability gaps may matter more on extended tasks than on short exchanges: an error early in a sequence can shape later actions. It also flags output volume as a practical concern. If an agent produces substantially more tokens while completing comparable work, cost and latency can accumulate across steps, even when the advertised per-token rate appears manageable.

The article also reports that its author observed confident false statements in hands-on use. That is an attributed personal observation, not a published benchmark finding. For an agent, an unsupported claim can become an input to later decisions, so buyers should test factual reliability and require suitable checks before assigning consequential tasks.

Amazon

AI model API key management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Large Jump From Mistral’s Earlier Scores

The same Artificial Analysis index version gave Mistral Large 3 a score of 9 and Medium 3.5 a score of 14, according to the source. Large 4’s 38.4 therefore represents a substantial advance on those earlier results. The source describes it as an unusually large step for a European lab, while noting that the gain does not put it among the top-scoring US or Chinese models.

The source’s comparisons also show why the launch framing needs qualification. Large 4 scores above some earlier Chinese models, including GLM-5.2 at 33.7 and DeepSeek V4 Pro at 36.0, but below newer entries such as GLM-5.3 and DeepSeek V4.1 Flash. Artificial Analysis figures are benchmark results, not a universal measure of model quality; performance depends on the task, setup and evaluation method.

“Reinforcement learning is still running.”

— Mistral, as reported by ThorstenMeyerAI.com

Amazon

large token capacity AI chatbot

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Reliability Remain Open

The supplied source says the model weights were promised for the end of October, but does not provide a calendar date for the release announcement or confirm that the weights were subsequently published. It also says the licence was unpublished at the time of writing. Until those details are available, developers cannot assess the terms for self-hosting or reuse from the source material alone.

Benchmark scores may shift because Mistral said reinforcement learning was continuing. The source does not provide a full account of its hands-on test method, sample size or the specific prompts behind the reported hallucinations. Nor does the supplied material establish how Large 4 performs across independent production deployments. Buyers should treat the author’s observation and the benchmark data as different kinds of evidence.

Amazon

AI model cost management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch for Weights and Updated Evaluations

The next milestones are publication of the model weights and licence, along with any updated Artificial Analysis results after further reinforcement learning. Those developments would clarify whether the model can be run outside Mistral’s API and whether its benchmark standing changes.

For now, teams considering Large 4 can compare it on their own agent tasks against alternatives, tracking completion rates, factual errors, token consumption, latency and total cost. The supplied source supports neither a blanket recommendation nor a conclusion that the model is unsuitable for every use. Its evidence instead points to a clear need for task-specific testing before production adoption.

Amazon

AI agent performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is a multimodal model available as a research preview through Mistral’s API. The source describes it as accepting text and images, producing text, and supporting a 512,000-token context window.

How did it score against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, as cited by the source. Several listed US and Chinese models scored higher; results are specific to this index version and should not be treated as a complete measure of performance on every task.

Is Large 4 open source?

Not according to the supplied source at the time of writing. It describes Large 4 as a proprietary API preview, with weights promised for the end of October and the licence not yet published.

What should agent developers test before using it?

Teams should measure task completion, factual reliability, output-token use, latency and total cost on their own workflows. The source reports high benchmark token use and an author-observed hallucination issue, but does not establish performance across all deployments.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Transform Proactive Cybersecurity For Large-Scale Organizations?

Google launches Fairwind, an AI-powered cybersecurity initiative offering rapid vulnerability patching for select organizations, raising questions about effectiveness and risks.

Grand Theft Auto 6 Leaks Response

Rockstar Games has issued a statement following the recent massive leak of Grand Theft Auto 6 gameplay footage and assets, confirming they are investigating the breach.

The Secret Behind Anthropic’s Advanced AI Watermark Technology

Anthropic has quietly deployed a sophisticated watermark in Claude’s responses, setting it apart from competitors and raising questions about detection reliability and future regulation.

10 Professional Mobile Workstation Laptops To Consider For AI In 2027

A sourced roundup of 10 mobile workstation configurations, with listed GPUs, memory and storage—and key details buyers should verify for AI work.