Skip to content

Memory budget · 12 GB

Best local LLMs for 12GB

The Nvidia GeForce RTX 4070 (12GB) gives about ~11 GB of its 12GB to model weights after the driver and display take their share. 76 of 155 local LLMs fit, biggest first.

Usable range
11 GB
Models that fit
76
Memory types
1
Top pick
14B

What 12GB actually gives you

Usable figures are sourced per device (tap a card for the full profile). Verdicts below use Q4_K_M, the community-default quant.

Every tracked 12GB device

Each device has its own profile with the full ranked list. Memory bandwidth, where listed, sets how fast tokens generate; it does not change what fits.

Top pick for 12GB Q4_K_M

Runs comfortably on the most capable 12GB setup (Nvidia GeForce RTX 4070 (12GB), ~11 GB usable) at ~9.4 GB. Check it against your exact device on its model page.

Models ranked for 12GB

Biggest that fits first GPU

Each chip links to the full breakdown for that model on a real 12GB device. "Tight" means it fits but with little headroom, close other apps.

The ceiling, per memory type

Nvidia GeForce RTX 4070 (12GB) (~11 GB usable)

Runs up to Ministral 3 14B (14B) comfortably at Q4_K_M. Larger models either sit tight or spill past the ~11 GB it can give a model.

Phones report 12GB too, but iOS/Android reserve more and the runtimes differ. Their usable pool is smaller:

DeepSeek-V2-Lite 16B gpt-oss 20B 21B ERNIE 4.5 21B-A3B 21B Mistral Small 3 24B 24B Sarvam-M 24B 24B Mistral Small 3.1 24B 24B Magistral Small 24B Devstral Small 24B LFM2 24B-A2B 24B Devstral Small 2 24B 24B Gemma 4 26B-A4B 26.5B Gemma 2 27B 27B Gemma 3 27B 27B Qwen3.8 27B 27B Qwen3.6 27B 27.8B Granite 4.1 30B 28.9B Sarvam-30B 30B Nemotron 3 Nano 30B-A3B 30B Nemotron Cascade 2 30B-A3B 30B GLM-4.7-Flash 30B Muse Glimmer 30B 30B Nemotron 3.5 Lightning 30B-A3B 30B Granite 4.2 30B 30B Qwen3 30B-A3B 30.5B Qwen3-Coder 30B-A3B 30.5B North Mini Code 1.0 30.5B Qwen3-VL 30B-A3B 31.1B Qwen2.5 32B 32B Qwen3 32B 32B DeepSeek-R1-Distill-Qwen 32B 32B Qwen2.5 Coder 32B 32B Granite 4.0 H Small 32B GLM-4-32B-0414 32B EXAONE 4.0 32B 32B OLMo 2 32B Instruct 32B Olmo 3.1 32B Instruct 32B Gemma 4 31B 32.7B Laguna XS 2.1 33.4B Yi 1.5 34B 34B Falcon-H1-34B-Instruct 34B Qwen-AgentWorld 35B-A3B 34.7B Command R 35B 35B Ornith 1.0 35B 35B Ornith 1.5 35B-A3B 35B Seed-OSS 36B Instruct 36B Qwen3.6 35B-A3B 36B Mixtral 8x7B 46.7B Llama-3.3-Nemotron-Super-49B-v1 49B Llama 3.3 70B 70B Qwen2.5 72B 72B Hunyuan-A13B-Instruct 80B Qwen3-Coder-Next 80B-A3B 80B Sarvam-105B 105B GLM-4.5-Air 106B Llama 4 Scout 109B Command A 111B gpt-oss 120B 117B Laguna S 2.1 118B Nemotron 3 Super 120B-A12B 120B Devstral 2 123B 123B dots.llm1 142B Qwen3 235B A22B 235B DeepSeek-V4-Flash 284B GLM-5.3-Flash 320B GLM-4.6 357B Llama 4 Maverick 400B MiniMax M3 428B MiniMax-M1-80k 456B Qwen3-Coder 480B-A35B Instruct 480B Nemotron 3 Ultra 550B-A55B 550B DeepSeek R1 671B DeepSeek V3 671B DeepSeek-R1-0528 671B GLM-5.2 744B Kimi K2 Instruct 1000B Kimi K2.6 1000B Kimi K2.7 Code 1000B DeepSeek-V4-Pro 1600B Kimi K3 2800B

FAQ

How much of 12GB can a model actually use?

It depends on the memory type. GPU VRAM: about 11 GB (Nvidia GeForce RTX 4070 (12GB)). The rest is reserved for the OS, display and runtime overhead.

What is the best local LLM for 12GB?

Ministral 3 14B (14B) is the strongest model that runs comfortably at Q4_K_M on the most capable 12GB setup (Nvidia GeForce RTX 4070 (12GB), ~11 GB usable). On a tighter 12GB device the ceiling is lower, shown per row above.

Sources

Memory figures are estimates at Q4_K_M with a small context. See methodology.