Model family · 5 sizes
Nemotron 3: which size runs locally?
Nemotron 3 comes in 5 sizes, from 4B to 550B. Some sizes are Mixture-of-Experts, so they run faster than their memory footprint suggests. Here is each size with its Q4_K_M weight, the memory it needs, and the hardware that runs it.
- Sizes
- 5
- Smallest
- 4B
- Largest
- 550B
- Runs from
- 8GB
The Nemotron 3 lineup
"Needs" is the sourced minimum memory for Q4_K_M with a small context. Larger context needs more.
Which Nemotron 3 fits your memory
Largest that fits: Nemotron 3 Nano 4B (4B), best case on Apple M1 (8GB).
Largest that fits: Nemotron 3 Nano 4B (4B), best case on Nvidia GeForce RTX 4080 (16GB).
Largest that fits: Nemotron 3 Nano 4B (4B), best case on Nvidia GeForce RTX 4090 (24GB).
Largest that fits: Nemotron Cascade 2 30B-A3B (30B), best case on Nvidia GeForce RTX 5090 (32GB).
Largest that fits: Nemotron Cascade 2 30B-A3B (30B), best case on Apple M5 Pro (48GB).
Largest that fits: Nemotron Cascade 2 30B-A3B (30B), best case on Apple M4 Max (64GB).
Largest that fits: Nemotron 3 Super 120B-A12B (120B), best case on Apple M5 Max (128GB).
Largest that fits: Nemotron 3 Super 120B-A12B (120B), best case on Apple M3 Ultra (256GB).
Best case means the most capable device at that size (usually a discrete GPU). A Mac at the same size sits roughly one rung lower; see the per-size breakdown on each memory budget page.
FAQ
Which Nemotron 3 size should I run locally?
Pick the largest size your memory allows. On 8GB (best case) up to Nemotron 3 Nano 4B; On 16GB (best case) up to Nemotron 3 Nano 4B; On 24GB (best case) up to Nemotron 3 Nano 4B; On 32GB (best case) up to Nemotron Cascade 2 30B-A3B; On 48GB (best case) up to Nemotron Cascade 2 30B-A3B; On 64GB (best case) up to Nemotron Cascade 2 30B-A3B; On 128GB (best case) up to Nemotron 3 Super 120B-A12B; On 256GB (best case) up to Nemotron 3 Super 120B-A12B. Smaller sizes run faster and leave headroom for context.
What is the smallest Nemotron 3 model?
Nemotron 3 Nano 4B at 4B parameters, about 2.64 GB on disk at Q4_K_M. It is the one to use on phones and 8 GB machines.
What is the largest Nemotron 3 model and what does it need?
Nemotron 3 Ultra 550B-A55B at 550B (mixture of experts), about 334.55 GB at Q4_K_M. It needs more than a typical 32 GB desktop; a high-memory Mac or multi-GPU rig.
Understand the numbers
Short guides to the ideas behind Nemotron 3's memory and quant figures.
Sources
- huggingface.co/bartowski/nvidia_Nemotron-3-Nano-30B-A3B-GGUF
- huggingface.co/bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF
- huggingface.co/mradermacher
- huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B
- huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-GGUF
- huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- huggingface.co/unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF
- huggingface.co/unsloth/NVIDIA-Nemotron-3-Ultra-550B-A55B-GGUF
- ollama.com/library/nemotron-3-nano
- ollama.com/library/nemotron-3-super
- ollama.com/library/nemotron-3-ultra
- ollama.com/library/nemotron-cascade-2
Memory figures are estimates at Q4_K_M. See methodology.