Model family · 14 sizes
Gemma: which size runs locally?
Gemma comes in 14 sizes, from 0.27B to 32.7B. Its strongest tracked here, Gemma 3 27B, scores an LMArena Elo of 1366. Here is each size with its Q4_K_M weight, the memory it needs, and the hardware that runs it.
- Sizes
- 14
- Smallest
- 0.27B
- Largest
- 32.7B
- Runs from
- 8GB
The Gemma lineup
- Gemma 3 270M0.27B
- Gemma 3 1B1B · ~0.81 GB Q4_K_M · needs ~2 GB
- Gemma 2 2B2.61B · ~1.71 GB Q4_K_M · needs ~3 GB
- Gemma 3 4B4B · ~2.49 GB Q4_K_M · needs ~4 GB · Elo 1303
- Gemma 4 E2B5.1B · ~3.11 GB Q4_K_M
- Gemma 3n E4B8B · ~4.23 GB Q4_K_M
- Gemma 4 E4B8B · ~4.98 GB Q4_K_M
- Gemma 2 9B9B · ~5.76 GB Q4_K_M · needs ~8 GB · Elo 1266
- Gemma 3 12B12B · ~7.3 GB Q4_K_M · needs ~10 GB · Elo 1342
- Gemma 4 12B12B · ~7.12 GB Q4_K_M
- Gemma 4 26B-A4B26.5B MoE · ~17.04 GB Q4_K_M
- Gemma 2 27B27B · ~16.65 GB Q4_K_M · needs ~20 GB · Elo 1289
- Gemma 3 27B27B · ~16.55 GB Q4_K_M · needs ~20 GB · Elo 1366
- Gemma 4 31B32.7B · ~18.32 GB Q4_K_M
"Needs" is the sourced minimum memory for Q4_K_M with a small context. Larger context needs more.
Which Gemma fits your memory
Largest that fits: Gemma 4 E2B (5.1B), best case on Apple M1 (8GB).
Largest that fits: Gemma 4 12B (12B), best case on Nvidia GeForce RTX 4080 (16GB).
Largest that fits: Gemma 4 31B (32.7B), best case on Nvidia GeForce RTX 4090 (24GB).
Largest that fits: Gemma 4 31B (32.7B), best case on Nvidia GeForce RTX 5090 (32GB).
Largest that fits: Gemma 4 31B (32.7B), best case on Apple M5 Pro (48GB).
Largest that fits: Gemma 4 31B (32.7B), best case on Apple M4 Max (64GB).
Largest that fits: Gemma 4 31B (32.7B), best case on Apple M5 Max (128GB).
Largest that fits: Gemma 4 31B (32.7B), best case on Apple M3 Ultra (256GB).
Best case means the most capable device at that size (usually a discrete GPU). A Mac at the same size sits roughly one rung lower; see the per-size breakdown on each memory budget page.
FAQ
Which Gemma size should I run locally?
Pick the largest size your memory allows. On 8GB (best case) up to Gemma 4 E2B; On 16GB (best case) up to Gemma 4 12B; On 24GB (best case) up to Gemma 4 31B; On 32GB (best case) up to Gemma 4 31B; On 48GB (best case) up to Gemma 4 31B; On 64GB (best case) up to Gemma 4 31B; On 128GB (best case) up to Gemma 4 31B; On 256GB (best case) up to Gemma 4 31B. Smaller sizes run faster and leave headroom for context.
What is the smallest Gemma model?
Gemma 3 270M at 0.27B parameters. It is the one to use on phones and 8 GB machines.
What is the largest Gemma model and what does it need?
Gemma 4 31B at 32.7B, about 18.32 GB at Q4_K_M. It fits a high-memory desktop GPU or Mac.
Understand the numbers
Short guides to the ideas behind Gemma's memory and quant figures.
Sources
- developers.googleblog.com
- gorilla.cs.berkeley.edu
- huggingface.co/bartowski/gemma-2-2b-it-GGUF
- huggingface.co/bartowski/google_gemma-3-1b-it-GGUF
- huggingface.co/bartowski/google_gemma-3-4b-it-GGUF
- huggingface.co/blog/gemma-july-update
- huggingface.co/blog/gemma2
- huggingface.co/ggml-org
- huggingface.co/google
- llm-stats.com
- lmarena.ai
- ollama.com/library/gemma2:2b
- ollama.com/library/gemma3
- ollama.com/library/gemma3:1b
- ollama.com/library/gemma3:270m
- ollama.com/library/gemma3/tags
Memory figures are estimates at Q4_K_M. See methodology.