Mistral AI: Mistral Small 4
mistralai/mistral-small-2603
Access Mistral Small 4 from Mistral AI using the Puter.js AI API.
Get Started// npm install @heyputer/puter.js
import { puter } from '@heyputer/puter.js';
puter.ai.chat("Explain quantum computing in simple terms", {
model: "mistralai/mistral-small-2603"
}).then(response => {
document.body.innerHTML = response.message.content;
});
<html>
<body>
<script src="https://js.puter.com/v2/"></script>
<script>
puter.ai.chat("Explain quantum computing in simple terms", {
model: "mistralai/mistral-small-2603"
}).then(response => {
document.body.innerHTML = response.message.content;
});
</script>
</body>
</html>
Model Card
Mistral Small 4 is a 119B-parameter open-source Mixture-of-Experts model (6B active per token) released under Apache 2.0, unifying instruction-following, reasoning, multimodal (text + image), and agentic coding into a single deployment. It features 128 experts, a 256k context window, and configurable reasoning effort that lets developers toggle between fast responses and deep step-by-step reasoning per request. Compared to its predecessor Mistral Small 3, it delivers 40% lower latency and 3x higher throughput while matching or surpassing GPT-OSS 120B on key benchmarks.
Context Window 256K
tokens
Max Output 256K
tokens
Input Cost $0.15
per million tokens
Output Cost $0.6
per million tokens
Input text, image
modalities
Tool Use Yes
Knowledge Cutoff Jun 2025
Release Date Mar 16, 2026
Output Speed 178
tokens / sec
Latency 0.56s
time to first token
Model Playground
Try Mistral Small 4 instantly in your browser.
This playground uses the Puter.js AI API — no API keys or setup required.
Benchmarks
How Mistral Small 4 performs on standard evaluations.
| Benchmark | Score |
|---|---|
| GPQA Diamond Graduate-level science Q&A | 76.9% |
| Humanity's Last Exam Cross-domain reasoning | 9.9% |
| SciCode Scientific programming | 38.8% |
| IFBench Instruction following | 48.2% |
| LCR Long-context reasoning | 49.7% |
| Terminal-Bench Hard Agentic terminal tasks | 17.4% |
| τ²-Bench Tool use / agents | 41.2% |
Scores sourced from Artificial Analysis.
Find other Mistral AI models →
Mistral Medium 3.5
Mistral Medium 3.5 is a dense 128-billion-parameter multimodal model from Mistral AI that unifies instruction-following, reasoning, and coding into a single set of weights. It features a 256k-token context window, native function calling, structured JSON output, and vision capabilities via a custom-trained encoder that handles variable image sizes. A per-request reasoning_effort parameter lets you toggle between fast responses and deeper chain-of-thought processing, making the same model suitable for quick chat replies and complex agentic workflows. On benchmarks, it scores 77.6% on SWE-Bench Verified and 91.4% on τ³-Telecom. It replaces Mistral's previous Medium 3.1, Magistral, and Devstral 2 models. Priced at $1.50 per million input tokens and $7.50 per million output tokens, it's a strong fit for developers building tool-calling agents, long-horizon coding tasks, and multi-step automation pipelines.
ChatMistral Medium 3.5
Mistral Medium 3.5 is a dense 128-billion-parameter multimodal model from Mistral AI that unifies instruction-following, reasoning, and coding into a single set of weights. This entry is Mistral's own dated direct-integration id for the same release available as mistralai/mistral-medium-3-5 through OpenRouter. It features a 256k-token context window, native function calling, structured JSON output, and vision capabilities via a custom-trained encoder that handles variable image sizes. A per-request reasoning_effort parameter lets you toggle between fast responses and deeper chain-of-thought processing, making the same model suitable for quick chat replies and complex agentic workflows. On benchmarks, it scores 77.6% on SWE-Bench Verified and 91.4% on τ³-Telecom. It replaces Mistral's previous Medium 3.1, Magistral, and Devstral 2 models. Priced at $1.50 per million input tokens and $7.50 per million output tokens, it's a strong fit for developers building tool-calling agents, long-horizon coding tasks, and multi-step automation pipelines.
ChatMinistral 14B
Ministral 14B is part of the Ministral 3 family, a 14B parameter multimodal model with vision capabilities under Apache 2.0. It offers advanced capabilities for local deployment with instruct, base, and reasoning variants achieving 85% on AIME'25.
Frequently Asked Questions
You can access Mistral Small 4 by Mistral AI through Puter.js AI API. Include the library in your web app or Node.js project and start making calls with just a few lines of JavaScript — no backend and no configuration required. You can also use it with Python or cURL via Puter's OpenAI-compatible API.
Mistral Small 4 is free to integrate using the Puter.js AI API. With the User-Pays Model, you can add AI to your app for $0, since users cover their own AI usage through their Puter account.
| Price per 1M tokens | |
|---|---|
| Input | $0.15 |
| Output | $0.6 |
Mistral Small 4 was created by Mistral AI and released on Mar 16, 2026.
Mistral Small 4 supports a context window of 256K tokens. For reference, that is roughly equivalent to 512 pages of text.
Mistral Small 4 can generate up to 256K tokens in a single response.
Mistral Small 4 has a knowledge cutoff date of Jun 2025. This means the model was trained on data available up to that date.
Mistral Small 4 accepts the following input types: text, image. It produces: text.
Yes, Mistral Small 4 supports tool use (function calling), allowing it to interact with external tools, APIs, and data sources as part of its response flow.
Mistral Small 4 scores 11.5 on the Artificial Analysis Intelligence Index, outperforming 49% of tracked models. On coding, it scores 26.6 (outperforms 33% of models).
Yes — the Mistral Small 4 API works with any JavaScript framework, Node.js, or plain HTML through Puter.js. Just include the library and start building. See the documentation for more details.
Add Mistral Small 4 to your app for free
Developers can integrate Mistral Small 4 for free using the Puter.js AI API.
With the User-Pays Model, each user covers their own AI usage instead of the developer.