Model family · 12 sizes
Mistral: which size runs locally?
Mistral comes in 12 sizes, from 3B to 123B. Its strongest tracked here, Mistral Small 3 24B, scores an LMArena Elo of 1357. Here is each size with its Q4_K_M weight, the memory it needs, and the hardware that runs it.
- Sizes
- 12
- Smallest
- 3B
- Largest
- 123B
- Runs from
- 8GB
The Mistral lineup
- Ministral 3 3B3B · ~2 GB Q4_K_M
- Mistral 7B7B · ~4.37 GB Q4_K_M · needs ~8 GB · Elo 1149
- Ministral 3 8B8B · ~4.84 GB Q4_K_M
- Mistral Nemo 12B12.2B · ~6.96 GB Q4_K_M · needs ~9 GB
- Ministral 3 14B14B · ~7.67 GB Q4_K_M
- Mistral Small 3 24B24B · ~14.33 GB Q4_K_M · needs ~16 GB · Elo 1357
- Mistral Small 3.1 24B24B · ~13.35 GB Q4_K_M
- Magistral Small24B · ~14 GB Q4_K_M
- Devstral Small24B · ~13.35 GB Q4_K_M
- Devstral Small 2 24B24B · ~13.35 GB Q4_K_M
- Mixtral 8x7B46.7B MoE · ~26.49 GB Q4_K_M · needs ~30 GB
- Devstral 2 123B123B · ~69.75 GB Q4_K_M
"Needs" is the sourced minimum memory for Q4_K_M with a small context. Larger context needs more.
Which Mistral fits your memory
Largest that fits: Ministral 3 8B (8B), best case on Nvidia GeForce RTX 3060 Ti (8GB). Comfortable up to Mistral 7B (7B).
Largest that fits: Ministral 3 14B (14B), best case on Nvidia GeForce RTX 4080 (16GB).
Largest that fits: Devstral Small 2 24B (24B), best case on Nvidia GeForce RTX 4090 (24GB).
Largest that fits: Mixtral 8x7B (46.7B), best case on Nvidia GeForce RTX 5090 (32GB). Comfortable up to Devstral Small 2 24B (24B).
Largest that fits: Mixtral 8x7B (46.7B), best case on Apple M5 Pro (48GB). Comfortable up to Devstral Small 2 24B (24B).
Largest that fits: Mixtral 8x7B (46.7B), best case on Apple M4 Max (64GB).
Largest that fits: Devstral 2 123B (123B), best case on Apple M5 Max (128GB).
Largest that fits: Devstral 2 123B (123B), best case on Apple M3 Ultra (256GB).
Best case means the most capable device at that size (usually a discrete GPU). A Mac at the same size sits roughly one rung lower; see the per-size breakdown on each memory budget page.
FAQ
Which Mistral size should I run locally?
Pick the largest size your memory allows. On 8GB (best case) up to Ministral 3 8B; On 16GB (best case) up to Ministral 3 14B; On 24GB (best case) up to Devstral Small 2 24B; On 32GB (best case) up to Mixtral 8x7B; On 48GB (best case) up to Mixtral 8x7B; On 64GB (best case) up to Mixtral 8x7B; On 128GB (best case) up to Devstral 2 123B; On 256GB (best case) up to Devstral 2 123B. Smaller sizes run faster and leave headroom for context.
What is the smallest Mistral model?
Ministral 3 3B at 3B parameters, about 2 GB on disk at Q4_K_M. It is the one to use on phones and 8 GB machines.
What is the largest Mistral model and what does it need?
Devstral 2 123B at 123B, about 69.75 GB at Q4_K_M. It fits a high-memory desktop GPU or Mac.
Understand the numbers
Short guides to the ideas behind Mistral's memory and quant figures.
Sources
- huggingface.co/bartowski/Mistral-7B-Instruct-v0.3-GGUF
- huggingface.co/bartowski/Mistral-Nemo-Instruct-2407-GGUF
- huggingface.co/mistralai/Ministral-3-14B-Instruct-2512
- huggingface.co/mistralai/Ministral-3-14B-Instruct-2512-GGUF
- huggingface.co/mistralai/Ministral-3-3B-Instruct-2512
- huggingface.co/mistralai/Ministral-3-3B-Instruct-2512-GGUF
- huggingface.co/mistralai/Ministral-3-8B-Instruct-2512
- huggingface.co/mistralai/Ministral-3-8B-Instruct-2512-GGUF
- huggingface.co/mistralai/Mistral-Nemo-Instruct-2407
- lmarena.ai
- ollama.com/library/ministral-3
- ollama.com/library/mistral
- ollama.com/library/mistral-nemo
- ollama.com/library/mistral-nemo/tags
- ollama.com/library/mistral-small3.2
- ollama.com/library/mistral/tags
Memory figures are estimates at Q4_K_M. See methodology.