Model family · 3 sizes
Kimi: which size runs locally?
Kimi comes in 3 sizes, from 1000B. Some sizes are Mixture-of-Experts, so they run faster than their memory footprint suggests. Here is each size with its Q4_K_M weight, the memory it needs, and the hardware that runs it.
- Sizes
- 3
- Smallest
- 1000B
- Largest
- 1000B
The Kimi lineup
"Needs" is the sourced minimum memory for Q4_K_M with a small context. Larger context needs more.
Which Kimi fits your memory
No Kimi size fits 8GB; even Kimi K2 Instruct needs more.
No Kimi size fits 16GB; even Kimi K2 Instruct needs more.
No Kimi size fits 24GB; even Kimi K2 Instruct needs more.
No Kimi size fits 32GB; even Kimi K2 Instruct needs more.
No Kimi size fits 48GB; even Kimi K2 Instruct needs more.
No Kimi size fits 64GB; even Kimi K2 Instruct needs more.
No Kimi size fits 128GB; even Kimi K2 Instruct needs more.
No Kimi size fits 256GB; even Kimi K2 Instruct needs more.
Best case means the most capable device at that size (usually a discrete GPU). A Mac at the same size sits roughly one rung lower; see the per-size breakdown on each memory budget page.
FAQ
Which Kimi size should I run locally?
Pick the largest size your memory allows. . Smaller sizes run faster and leave headroom for context.
What is the smallest Kimi model?
Kimi K2 Instruct at 1000B parameters, about 578.15 GB on disk at Q4_K_M. It is the one to use on phones and 8 GB machines.
What is the largest Kimi model and what does it need?
Kimi K2.7 Code at 1000B (mixture of experts), about 583.71 GB at Q4_K_M. It needs more than a typical 32 GB desktop; a high-memory Mac or multi-GPU rig.
Understand the numbers
Short guides to the ideas behind Kimi's memory and quant figures.
Sources
- aider.chat
- gorilla.cs.berkeley.edu
- hpcwire.com
- huggingface.co/moonshotai/Kimi-K2-Instruct
- huggingface.co/moonshotai/Kimi-K2.6
- huggingface.co/moonshotai/Kimi-K2.7-Code
- huggingface.co/unsloth/Kimi-K2-Instruct-GGUF
- huggingface.co/unsloth/Kimi-K2.6-GGUF
- huggingface.co/unsloth/Kimi-K2.7-Code-GGUF
- ollama.com
Memory figures are estimates at Q4_K_M. See methodology.