What is Mistral and how it can help you be more productive

Mistral AI assistant guide · Updated September 2026

Mistral isn’t just “Europe’s open-source ChatGPT” anymore. What started in 2023 as a Paris-based lab founded by former Meta and DeepMind researchers has grown into a full open-weight platform: efficient frontier-adjacent models, a unified agent called Vibe (formerly Le Chat), strong coding tools, document AI, and speech models – all with a heavy emphasis on open weights, European data residency, and aggressive pricing. In April 2026 Mistral shipped Medium 3.5, its dense flagship that folds chat, reasoning, coding, and multimodal capabilities into one model, while Large 3 (December 2025) remains the standout open-weight MoE.

In this guide we explain what Mistral is in 2026, how the Medium 3.5 / Large 3 / Small 4 / Ministral 3 family works, what Vibe, coding agents, OCR, and open weights can do, what the plans cost, and how to use it well.

Sep 2023First public release, Mistral 7B via torrent
Medium 3.5 · Large 3Current flagship dense + open MoE
Vibe · Open weightsAgentic surface + self-hostable models
01 · What it is

What is Mistral?

Mistral is the family of AI models and products from Mistral AI, the European lab founded in April 2023 in Paris by Arthur Mensch (CEO, ex-Google DeepMind), Guillaume Lample (Chief Scientist, ex-Meta), and Timothée Lacroix (CTO, ex-Meta). The name nods to the strong Mediterranean wind – a deliberate choice for a company that wants to move fast and stay independent. Mistral quickly became known for releasing capable open-weight models under permissive licenses (mostly Apache 2.0 or Modified MIT) while also running a commercial API and consumer assistant.

Unlike OpenAI, Anthropic, or Google, Mistral’s defining trait is openness and portability. Most of its models ship with downloadable weights so enterprises can inspect, fine-tune, and self-host the intelligence they build on. That, plus EU data residency and in-region inference, is why Mistral has become the default choice for European institutions chasing digital sovereignty. By September 2026 it had passed 1,000 employees and a €12B valuation, with an 11% stake from ASML.

Today, depending on the surface and plan, Mistral can:

  • Answer questions & explain concepts
  • Write, rewrite, summarize & translate
  • Search the web and ground answers
  • Reason with adjustable effort
  • Analyze files, documents, images & data
  • Write, debug & ship code
  • Run as an agent across tools and long tasks
  • Handle OCR and document understanding at high speed
  • Generate speech (TTS) and transcribe audio
  • Run locally or self-host thanks to open weights
Mistral – the models

Mistral’s current generation centers on Medium 3.5 (dense 128B flagship for agentic/coding work), Large 3 (open-weight MoE with 41B active / 675B total parameters), Small 4 (efficient hybrid multimodal MoE), and the Ministral 3 edge family (3B/8B/14B). Specialized lines include Codestral (code completion), OCR 4.1, and Voxtral (speech).

Mistral – the products

You meet Mistral through Vibe (the renamed Le Chat – web, mobile, CLI, and VS Code extension), the Mistral API / Studio (developers), and downloadable open weights on Hugging Face. The same family powers all of them, with EU data residency and sovereign compute options.

02 · The model family

Medium 3.5, Large 3, Small 4, and Ministral 3

Mistral’s lineup in September 2026 emphasizes fewer, more capable models that absorb what used to be separate specialist lines (Pixtral for vision, Magistral for reasoning, Devstral for coding). Le Chat became Vibe on May 28, 2026, unifying work and code under one agent.

Frontier agentic · dense · Apr 28, 2026

Medium 3.5

128B dense parameters, 256K context, multimodal, with adjustable reasoning effort. Optimized for long-horizon agentic tasks, tool use, and coding. Open weights under Modified MIT. The practical workhorse for serious agent and coding sessions.

Maximum capability for agents & codeDense flagship
Open-weight flagship · Dec 2, 2025

Large 3

Sparse MoE (41B active / 675B total), multimodal, multilingual, Apache 2.0. One of the strongest permissive open-weight models; competitive on general instruction following while remaining relatively cheap to serve via API ($0.50 / $1.50 per million tokens).

Best open-weight general model
Efficient hybrid · Mar 16, 2026

Small 4

~119B total / ~6B active MoE, 256K context, native multimodal + configurable reasoning + coding in one model. Apache 2.0. Excellent cost/performance for everyday professional work.

Best value everyday model
Edge & local · Dec 2, 2025

Ministral 3

Compact multimodal models (3B/8B/14B) with vision, tool calling, and strong multilingual support. Symmetric low pricing and open weights make them practical for edge devices and high-volume workloads.

Lowest cost / on-device

Specialized companions include Codestral (high-quality code completion/FIM), OCR 4.1 (fast, high-accuracy document extraction with structural understanding, Aug 2026), and Voxtral (transcription + TTS with voice cloning). Medium 3.5 self-hosts on as few as four GPUs; Small 4 is Apache 2.0.

03 · How it works

How does Mistral work?

Mistral models are transformers (dense or sparse MoE). Medium 3.5 and Small 4 unify capabilities that used to require separate models, with a reasoning-effort control that lets you trade latency for deeper step-by-step thinking. Large 3 uses mixture-of-experts so only a fraction of parameters activate per token, keeping inference efficient despite the large total parameter count.

Modern Mistral combines the model with tools, retrieval, and connectors. A typical request flows through several stages:

  • 1

    Understand the request

    Reads the prompt, conversation history, uploaded files/images, and any connected tools, libraries, or memory.

  • 2

    Reason through the problem

    Adjustable effort (or built-in agentic planning in Vibe) decides how much internal deliberation to apply – from instant answers to deep multi-step reasoning.

  • 3

    Decide which tools to use

    Web search, code execution, OCR, document libraries, connectors, image generation, and Agentic Search that opens and verifies sources.

  • 4

    Generate the response or take the action

    Produces the answer, code, document, or completed multi-step task – sometimes after a remote agent has worked independently in the cloud.

  • 5

    Refine with follow-ups

    Canvas-style editing, project folders, and persistent agent sessions keep context alive. Agentic Search cuts turns and token use on FinanceBench and OfficeQA Pro.

04 · What you can do

What can you do with Mistral?

Research and information

Web search grounds answers in current results. Vibe agents run multi-step research and long-running workflows, with Agentic Search (Aug 20, 2026) that opens and verifies sources before answering.

Writing and editing

Draft, rewrite, summarize, and translate. Strong multilingual performance (especially European languages) and project/library features for recurring work. Canvas keeps documents editable.

Coding and DevOps

Medium 3.5 and the coding stack (Vibe CLI / VS Code extension, Codestral for completion) handle feature work, debugging, tests, and PR-style delivery. Open weights let you self-host or fine-tune on your own infra.

Document AI and OCR

OCR 4.1 extracts structured content (paragraphs, tables, equations) from PDFs and office files at high speed and accuracy, including on-prem options for regulated environments. $4–5 per 1,000 pages.

Data analysis

Upload spreadsheets or documents; the models clean, compute, and explain results, often via code execution, rendering charts and dashboards inside the conversation.

Multimodal and voice

Native image understanding across the main models, image generation, Voxtral transcription/TTS, and real-time voice modes. TTS from 3 sec of reference audio with voice cloning.

Self-hosting and control

Many flagship and mid-tier models ship open weights under permissive licenses. Enterprises can keep data in-house, fine-tune, or run on Mistral Compute with EU data residency.

Agents and integrations

Vibe acts as a unified agent for knowledge work and coding. Connectors, MCP-style tools, libraries, and task scheduling let it operate across apps and long-running jobs.

05 · From chatbot to agent

The agentic shift: Vibe

Mistral’s biggest product evolution is Vibe – the rebranded and expanded Le Chat that now serves both everyday knowledge work and serious coding in one surface (web, mobile, CLI, IDE extension). Give it a goal and it plans, uses tools, and returns finished work.

01

Vibe Work

Productivity mode for multi-step tasks. Give it a goal – research a topic, clean a folder of receipts, implement a feature – and it plans, uses tools, and returns finished work. Connectors are on by default, pulling in the context it needs. Free tier is usable for lighter tasks; Pro unlocks all-day coding and higher limits.

One agent for work and code · May 28, 2026 rename

02

Vibe Code

Developer mode via CLI, VS Code extension, or remote web sessions. It reads files, edits code, runs commands, and opens pull requests. Remote agents run in the cloud, notify you when done, and keep you in the loop on sensitive actions.

CLI · VS Code · Remote web sessions

03

Studio & Sovereign Compute

Studio builds agents and RAG pipelines with Workflows, Libraries, and Observability. Mistral Compute offers regional inference (EU/US) and European Compute Units (200 MW by end-2027, 1 GW by 2030). Sovereign option with on-prem, white-label, and custom models.

Studio · Forge · Mistral Compute

The philosophy is the same as rivals: move from “answer my question” to “finish this outcome,” while keeping strong open-weight options so you are not locked into a single vendor’s cloud. Agentic Search (Aug 20, 2026) now powers document AI with multi-step retrieval that verifies evidence before answering.

06 · Key moments

Mistral timeline: 2023–2026

Apr 2023

Mistral AI founded in Paris by Arthur Mensch, Guillaume Lample, and Timothée Lacroix; €105M seed follows in June.

Sep 2023

Mistral 7B released via BitTorrent magnet link under Apache 2.0 – first open-weight model, beating Llama 2 13B with 7B params.

Dec 2023

Mixtral 8x7B introduces sparse MoE (46.7B usable, 12.9B active per token); €385M Series A at €2B valuation.

Feb 2024

Mistral Large and Small launch; Microsoft invests $16M; Large first on Azure. Le Chat launches as consumer chatbot.

Dec 2025

Mistral 3 family ships: Large 3 (675B/41B active, 256K, Apache 2.0), Ministral 3 (3/8/14B), Devstral 2, plus Codestral 25.08 refresh.

Mar 16, 2026

Small 4 ships as unified reasoning/vision/coding MoE (119B total / 6B active, 256K, Apache 2.0).

Apr 28, 2026

Medium 3.5 launches as dense agentic/coding flagship (128B, 256K, Modified MIT) – replaces Magistral and Devstral 2 in Vibe. Medium 3.1 refresh follows Aug 12.

May 28, 2026

Le Chat renamed Vibe; Work Mode and Code Mode introduced; Vibe CLI and VS Code extension launched. Emmi AI acquired; HUMAIN Saudi collaboration announced Aug 24.

Aug 2026

Regional endpoints, Priority Tier, ECUs, and 1 GW by 2030 compute plan (Aug 11); OCR 4.1 (Aug 13); Agentic Search (Aug 20); sovereign AI deals and Microsoft multibillion infra expansion.

07 · Staying productive

How to use Mistral well

Be specific

State audience, goal, format, and constraints up front. “Turn this into five board-ready bullets with one evidence line each” beats “summarize this.”

Match model to job

Everyday / high volume → Small 4 or Ministral. General quality + open weights → Large 3. Serious agentic or coding work → Medium 3.5. Code completion → Codestral. Documents → OCR 4.1.

Use Vibe’s agent mode

For multi-step outcomes, describe the finished deliverable and let the agent plan and execute rather than micromanaging every turn. Skills and project folders keep context alive.

The prompt pattern that works

Role / Context:You are a senior engineer reviewing this PR.Task:Identify bugs, style issues, and missing tests.Format:Bullet list grouped by severity, with file:line references.

Give clear role + task + format, attach the real files or codebase context, and iterate. For recurring work, use project folders or libraries so context does not have to be re-pasted. Leverage open weights when privacy or cost at scale matters.

08 · The window

How large is the context window?

Current main models offer 256K-token context windows – roughly 192,000 words – enough for long documents, substantial codebases, or multi-turn agent sessions. Ministral edge models range 131K–262K depending on size. The hosted GLM 5.2 extends to 1M for long-context workflows.

Medium 3.5 / Small 4 / Large 3 – 256K tokens≈ 192,000 words
Ministral 3 (3B/8B/14B) – 131K–262K tokens≈ 98K–196K words
Mistral 7B (2023) – 32K tokens≈ 24,000 words

Practical rule of thumb: a token is roughly 0.75 English words, and about 1.5 tokens per word once you factor in punctuation and spacing. As with every frontier model, attention is not perfectly uniform across the entire window.

09 · Plans and pricing

What does Mistral cost?

PlanPriceModel accessLimits
Free$0Medium 3.5 (limited)Limited messages, web searches, coding sessions; image generation; 5 scheduled tasks; EU residency
Pro$14.99/moMedium 3.5Up to 6x Free’s messages, 5x web searches, 40x image generations; all-day coding; 15GB libraries
Team$24.99/user/moMedium 3.5Up to 6x Free’s messages per user; 30GB per user; admin controls; domain verification
EnterpriseCustomCustom modelsCustom models, agents, workflows; audit logs; SAML SSO; white label; opt-out of training
API · Medium 3.5$1.50 / $7.50per 1M tokens in/outFlagship agentic/coding
API · Large 3$0.50 / $1.50per 1M tokens in/outStrong open-weight value
API · Small 4$0.15 / $0.60per 1M tokens in/outEfficient everyday
API · Ministral 3$0.10 – $0.20per 1M tokens in/outEdge / high volume, symmetric pricing
API · Codestral$0.30 / $0.90per 1M tokens in/outCode completion · batch -50% · cached -90%

Cached input is heavily discounted (often ~90% off). Batch processing can cut prices further. Open-weight models can also be self-hosted at pure hardware cost. The chart above reflects September 2026.

10 · FAQ

Frequently asked questions

Is Mistral free?

Yes – a capable free tier exists on Vibe, and many models are fully open-weight so you can run them yourself at no per-token cost. Pro at $14.99 is the cheapest paid tier from a major provider.

How does Mistral compare to ChatGPT, Claude, Gemini, or Grok?

Mistral usually sits just behind the absolute closed-frontier leaders on the hardest reasoning and agent benchmarks, but it is highly competitive on cost, open weights, European compliance/data residency, multilingual performance, coding, and document AI. For many production and self-hosted use cases the combination of quality + price + openness is the deciding factor.

Can I self-host the models?

Yes for Large 3, Small 4, Ministral 3, Medium 3.5 (Modified MIT), and most open releases. Quantized versions of the smaller models run on high-end consumer or modest server GPUs; Medium 3.5 self-hosts on as few as four GPUs.

Is my data used for training?

Paid Team and Enterprise can opt out of model training. On Free, data may be used to improve models. Self-hosting removes the question entirely – nothing leaves your infrastructure.

Does Mistral have an app?

Yes – Vibe on web at chat.mistral.ai, iOS and Android apps, Vibe CLI, VS Code extension, and remote agents that run in the cloud and notify you when they finish. Le Chat history and settings carried over.

Heads-up:Mistral is a fast-moving product, and the lab shipped its entire 3rd generation plus Voxtral and Devstral in the past 9 months alone. Plan names, model names, and prices change frequently – the details above reflect September 2026 (Medium 3.5 Apr 28, Large 3 Dec 2025, Vibe rename May 28). Check Mistral’s official La Plateforme and Le Chat pages for current availability and limits.

Related reading