Skip to content

Device profile · Windows

Best local LLMs for Nvidia GeForce RTX 2060 (6GB)

Nvidia GeForce RTX 2060 (6GB) has ~5 GB usable for model weights and runs 38 of 149 popular models. Best tool: LM Studio.

36 run well 2 tight fit 111 too heavy
Usable memory
~5 GB
Models run
38
Too large
111
Top pick
4B
Top pick Q4_K_M

Runs at Q4_K_M using ~3.8 GB of ~5 GB usable.

Runs on Nvidia GeForce RTX 2060 (6GB)

Compatible models 38 total
Kimi K3DeepSeek-V4-ProKimi K2 InstructKimi K2.6Kimi K2.7 CodeGLM-5.2DeepSeek R1DeepSeek V3DeepSeek-R1-0528Nemotron 3 Ultra 550B-A55BQwen3-Coder 480B-A35B InstructMiniMax-M1-80kMiniMax M3Llama 4 MaverickGLM-4.6GLM-5.3-FlashDeepSeek-V4-FlashQwen3 235B A22Bdots.llm1Devstral 2 123BNemotron 3 Super 120B-A12BLaguna S 2.1gpt-oss 120BCommand ALlama 4 ScoutGLM-4.5-AirSarvam-105BHunyuan-A13B-InstructQwen2.5 72BLlama 3.3 70BLlama-3.3-Nemotron-Super-49B-v1Mixtral 8x7BSeed-OSS 36B InstructQwen3.6 35B-A3BCommand R 35BOrnith 1.0 35BOrnith 1.5 35B-A3BQwen-AgentWorld 35B-A3BYi 1.5 34BFalcon-H1-34B-InstructLaguna XS 2.1Gemma 4 31BQwen2.5 32BQwen3 32BDeepSeek-R1-Distill-Qwen 32BQwen2.5 Coder 32BGranite 4.0 H SmallGLM-4-32B-0414EXAONE 4.0 32BOLMo 2 32B InstructGranite 4.0 H SmallOlmo 3.1 32B InstructQwen3 30B-A3BQwen3-Coder 30B-A3BNorth Mini Code 1.0Sarvam-30BNemotron 3 Nano 30B-A3BNemotron Cascade 2 30B-A3BGLM-4.7-FlashMuse Glimmer 30BNemotron 3.5 Lightning 30B-A3BGranite 4.1 30BQwen3.6 27BGemma 2 27BGemma 3 27BQwen3.8 27BGemma 4 26B-A4BMistral Small 3 24BSarvam-M 24BMistral Small 3.1 24BMagistral SmallDevstral SmallLFM2 24B-A2BDevstral Small 2 24Bgpt-oss 20BERNIE 4.5 21B-A3BDeepSeek-V2-LitePhi-4 14BQwen2.5 14BQwen3 14BDeepSeek-R1-Distill-Qwen 14BQwen2.5 Coder 14BPhi-4-reasoningMinistral 3 14BMistral Nemo 12BGemma 3 12BGemma 4 12BLlama 3.2 Vision 11BFalcon3 10BGemma 2 9BGLM-4 9BGLM-4-9B-0414Nemotron Nano 9B v2Ornith 1.0 9BOrnith 1.5 9BGranite 4.1 8BLFM2.5 8B-A1BQwen2.5-VL 7BDeepSeek-R1-0528-Qwen3-8BLlama 3.1 8BQwen3 8BDeepSeek-R1-Distill-Llama 8BGemma 3n E4BGemma 4 E4BMinistral 3 8BRNJ-1 8BMistral 7BQwen2.5 7BDeepSeek-R1-Distill-Qwen 7BQwen2.5 Coder 7BOlmo 3 7B Instruct

Best way to run models on Windows

Runtime guide Windows

Beginner: LM Studio, Best GUI on Windows, auto-detects CUDA/Vulkan backends.

Power user: Ollama (CUDA), Scriptable server; CUDA path is fastest on NVIDIA.

AMD GPUs run via Vulkan/ROCm at roughly half CUDA throughput. NVIDIA is the smooth path on Windows.

Full Windows tool guide →

FAQ

What is the best local LLM for Nvidia GeForce RTX 2060 (6GB)?

Gemma 3 4B is the strongest model that runs comfortably, using ~3.8 GB at Q4_K_M of the ~5 GB usable on Nvidia GeForce RTX 2060 (6GB).

How much of Nvidia GeForce RTX 2060 (6GB)'s memory can I use for a model?

About 5 GB. On a discrete GPU, leave ~1 GB of VRAM for the driver and display.

Which tool should I use on Windows?

LM Studio (Best GUI on Windows, auto-detects CUDA/Vulkan backends.) or Ollama (CUDA) for speed. AMD GPUs run via Vulkan/ROCm at roughly half CUDA throughput. NVIDIA is the smooth path on Windows.

Compare nearby devices

Hardware with a similar usable pool, and how many of the 149 tracked models each runs next to Nvidia GeForce RTX 2060 (6GB)'s 38.

Sources

Memory figures are estimates. See methodology.