What you’ll walk away with
By the end of this you’ll have a working Python call to Mistral Large 4 that takes text + image input, toggles reasoning on/off, and can return structured JSON – on the public preview that opened October 6, 2026. Weights can wait.
Mistral Large 4 (ML4, also called le Chonk) is a hybrid instruct+reasoning MoE: about 1.05T total parameters, 52B active, plus a 1.6B vision encoder. Text and images in, text out. Docs list a 1M-token context. European training run, 160+ languages. The API is live in Mistral Studio now; open weights are promised for the end of October 2026.
Quick setup: key + SDK in five minutes
Open console.mistral.ai, create a key under API Keys, export it:
export MISTRAL_API_KEY="your_key_here"
Install the client:
pip install mistralai
Base URL: https://api.mistral.ai/v1. Preview sale pricing (as of early October 2026): $0.68 / 1M input, $0.07 cached, $2.09 output – list is $1.36 / $0.14 / $4.18; batch is cheaper still (Mistral pricing docs).
API keys feel like pure ceremony until the wrong model string quietly ruins a whole afternoon. That ceremony is next.
First call and the model-ID trap
Hardcode mistral-large-4 or mistral-large-4-0. Do not use mistral-large-latest.
import os
from mistralai import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
response = client.chat.complete(
model="mistral-large-4",
messages=[{"role": "user", "content": "Explain MoE routing in two plain sentences."}],
temperature=0.2
)
print(response.choices[0].message.content)
print("Actual model:", response.model) # always verify
Turns out mistral-large-latest still resolved to Large 3 through the launch window. Print response.model every time. That single field ends the “why does this feel old?” loop.
Turn reasoning on or off + vision
Large 4 thinks unless you stop it. reasoning_effort is the dial – "high" emits a full thinking chunk (more tokens, more spend); "none" drops that chunk.
response = client.chat.complete(
model="mistral-large-4",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Locate the main connector in this schematic and describe its pinout."},
{"type": "image_url", "image_url": "https://example.com/schematic.png"}
]
}
],
reasoning_effort="high", # or "none"
response_format={"type": "json_object"} # optional structured
)
Image + text parts sit in the same content array. Function calling, structured outputs, and agents use the same patterns as the rest of the Mistral API (model card).
Default path for production:
reasoning_effort="none". Flip to high only for multi-step or high-stakes work. Thinking tokens are a silent bill multiplier on routine prompts.
Reasoning tokens behave like leaving a taxi meter running while the driver “thinks” about the route. Useful on a hard detour. Wasteful for a two-block hop – and independent checks (simple arithmetic at temperature 0) still saw dozens of think tokens until "none".
Under launch-week load, expect roughly 50 tokens/s and 15-20 s time-to-first-token on the preview; Artificial Analysis logged about 116 t/s and 18.7 s TTFT under their conditions, and Mistral eng posts acknowledged the spike. Bake in retries and timeouts now. This may ease after the first weeks.
Common pitfalls that burn tokens or break calls
- Wrong model string → older capability with no loud error. Trust
response.model, not the alias. - Leaving reasoning on → long outputs and surprise cost. Explicit
"none"is the safe default. - Assuming full 1M context on every gateway → docs say 1M; Artificial Analysis / OpenRouter-style listings show ~512k-524k. Oversize prompts can 400 or truncate on proxies. Measure your real max prompt.
- Ignoring launch-week limits → 429s on free/low tiers. Read rate-limit headers; batch bulk jobs.
- Ungrounded factual answers → 41.9% hallucination on AA-Omniscience knowledge probes. RAG or tool-ground anything customer-facing.
What the numbers look like right now
| Metric | Mistral Large 4 Preview | Notes |
|---|---|---|
| AA Intelligence Index | ~38.4 | As of early preview measurements |
| Cyber Index | 50 | Strong open-side showing on tasks closed models often refuse |
| DeepSWE v1.1 | 61.7% | Agentic coding |
| Output speed (AA) | ~116 t/s | Falls hard under load |
| Sale price (in/out) | $0.68 / $2.09 | Per 1M tokens; half list; early preview window |
Cyber, visual grounding, and long-horizon agent work are the sweet spots per Mistral’s October 2026 announcement and the independent runs above. Cheaper or faster frontier options exist; this isn’t them.
When you should skip Mistral Large 4
Skip it for low-latency chat widgets – TTFT hurts. Skip ungrounded trivia or support bots given the knowledge-probe hallucination rate. Skip if you need top coding Elo today or self-host this week; weights are still end-of-October 2026, and license/VRAM figures were unannounced at preview launch.
Pick it for European data residency, a future open-weight path, multimodal + agentic in one API, or cyber/enterprise jobs that closed models refuse. For cheap high-volume classification, smaller Mistral models on the same platform cost less.
FAQ
Is Mistral Large 4 free to try?
No. Preview API is pay-as-you-go at the sale rates above. Check the console for whatever trial credit your account shows.
Can I run it locally yet?
Not today. Hosted API (or gateways that proxy it) only until weights land at the end of October 2026. A 1T-class MoE will need serious hardware once checkpoints appear – exact VRAM and license terms were still open questions at launch.
Does reasoning_effort change quality or just formatting?
Both. People treat it like a free quality boost. It isn’t. "high" forces longer internal deliberation and returns a thinking chunk; hard multi-hop agent work often improves. Simple chat gets slower and pricier for almost no gain – basic arithmetic still burned dozens of thinking tokens in early tests until "none". Use it as a cost/behavior dial, not a default upgrade.
Open Studio, paste the first snippet, send a real image + text prompt to mistral-large-4 before sale pricing or load shifts. Fastest way to see if le Chonk belongs in your stack.