Skip to content

Memory budget · 6 GB

Best local LLMs for 6GB

The Nvidia GeForce RTX 2060 (6GB) gives about ~5 GB of its 6GB to model weights after the driver and display take their share. 40 of 155 local LLMs fit, biggest first.

Usable range
5 GB
Models that fit
40
Memory types
1
Top pick
4B

What 6GB actually gives you

Usable figures are sourced per device (tap a card for the full profile). Verdicts below use Q4_K_M, the community-default quant.

Top pick for 6GB Q4_K_M

Runs comfortably on the most capable 6GB setup (Nvidia GeForce RTX 2060 (6GB), ~5 GB usable) at ~3.8 GB. Check it against your exact device on its model page.

Models ranked for 6GB

Biggest that fits first GPU

Each chip links to the full breakdown for that model on a real 6GB device. "Tight" means it fits but with little headroom, close other apps.

The ceiling, per memory type

Nvidia GeForce RTX 2060 (6GB) (~5 GB usable)

Runs up to Nemotron 3 Nano 4B (4B) comfortably at Q4_K_M. Larger models either sit tight or spill past the ~5 GB it can give a model.

Mistral 7B 7B Qwen2.5 7B 7B DeepSeek-R1-Distill-Qwen 7B 7B Qwen2.5 Coder 7B 7B Olmo 3 7B Instruct 7B Llama 3.1 8B 8B Qwen3 8B 8B DeepSeek-R1-Distill-Llama 8B 8B Gemma 3n E4B 8B Gemma 4 E4B 8B Ministral 3 8B 8B RNJ-1 8B 8B Granite 4.2 8B 8B DeepSeek-R1-0528-Qwen3-8B 8.19B Qwen2.5-VL 7B 8.29B LFM2.5 8B-A1B 8.3B Qwen3-VL 8B 8.77B Granite 4.1 8B 8.8B Gemma 2 9B 9B GLM-4 9B 9B GLM-4-9B-0414 9B Nemotron Nano 9B v2 9B Ornith 1.0 9B 9B Ornith 1.5 9B 9B Falcon3 10B 10B Llama 3.2 Vision 11B 10.7B Gemma 3 12B 12B Gemma 4 12B 12B Mistral Nemo 12B 12.2B Phi-4 14B 14B Qwen2.5 14B 14B Qwen3 14B 14B DeepSeek-R1-Distill-Qwen 14B 14B Qwen2.5 Coder 14B 14B Phi-4-reasoning 14B Ministral 3 14B 14B DeepSeek-V2-Lite 16B gpt-oss 20B 21B ERNIE 4.5 21B-A3B 21B Mistral Small 3 24B 24B Sarvam-M 24B 24B Mistral Small 3.1 24B 24B Magistral Small 24B Devstral Small 24B LFM2 24B-A2B 24B Devstral Small 2 24B 24B Gemma 4 26B-A4B 26.5B Gemma 2 27B 27B Gemma 3 27B 27B Qwen3.8 27B 27B Qwen3.6 27B 27.8B Granite 4.1 30B 28.9B Sarvam-30B 30B Nemotron 3 Nano 30B-A3B 30B Nemotron Cascade 2 30B-A3B 30B GLM-4.7-Flash 30B Muse Glimmer 30B 30B Nemotron 3.5 Lightning 30B-A3B 30B Granite 4.2 30B 30B Qwen3 30B-A3B 30.5B Qwen3-Coder 30B-A3B 30.5B North Mini Code 1.0 30.5B Qwen3-VL 30B-A3B 31.1B Qwen2.5 32B 32B Qwen3 32B 32B DeepSeek-R1-Distill-Qwen 32B 32B Qwen2.5 Coder 32B 32B Granite 4.0 H Small 32B GLM-4-32B-0414 32B EXAONE 4.0 32B 32B OLMo 2 32B Instruct 32B Olmo 3.1 32B Instruct 32B Gemma 4 31B 32.7B Laguna XS 2.1 33.4B Yi 1.5 34B 34B Falcon-H1-34B-Instruct 34B Qwen-AgentWorld 35B-A3B 34.7B Command R 35B 35B Ornith 1.0 35B 35B Ornith 1.5 35B-A3B 35B Seed-OSS 36B Instruct 36B Qwen3.6 35B-A3B 36B Mixtral 8x7B 46.7B Llama-3.3-Nemotron-Super-49B-v1 49B Llama 3.3 70B 70B Qwen2.5 72B 72B Hunyuan-A13B-Instruct 80B Qwen3-Coder-Next 80B-A3B 80B Sarvam-105B 105B GLM-4.5-Air 106B Llama 4 Scout 109B Command A 111B gpt-oss 120B 117B Laguna S 2.1 118B Nemotron 3 Super 120B-A12B 120B Devstral 2 123B 123B dots.llm1 142B Qwen3 235B A22B 235B DeepSeek-V4-Flash 284B GLM-5.3-Flash 320B GLM-4.6 357B Llama 4 Maverick 400B MiniMax M3 428B MiniMax-M1-80k 456B Qwen3-Coder 480B-A35B Instruct 480B Nemotron 3 Ultra 550B-A55B 550B DeepSeek R1 671B DeepSeek V3 671B DeepSeek-R1-0528 671B GLM-5.2 744B Kimi K2 Instruct 1000B Kimi K2.6 1000B Kimi K2.7 Code 1000B DeepSeek-V4-Pro 1600B Kimi K3 2800B

FAQ

How much of 6GB can a model actually use?

It depends on the memory type. GPU VRAM: about 5 GB (Nvidia GeForce RTX 2060 (6GB)). The rest is reserved for the OS, display and runtime overhead.

What is the best local LLM for 6GB?

Gemma 3 4B (4B) is the strongest model that runs comfortably at Q4_K_M on the most capable 6GB setup (Nvidia GeForce RTX 2060 (6GB), ~5 GB usable). On a tighter 6GB device the ceiling is lower, shown per row above.

Sources

Memory figures are estimates at Q4_K_M with a small context. See methodology.