Leaderboard · LMArena
Best local LLMs by human preference
LMArena ranks models by Elo from blind, head-to-head human votes. It captures overall chat quality and instruction-following as people actually perceive it, which is why it is the most-cited general leaderboard.
- 1 1383 Qwen3 30B-A3B 30.5B MoE Apple M5 (32GB)~20.7 GB
- 2 1366 Gemma 3 27B 27B Apple M5 (32GB)~18.6 GB
- 3 1357 Mistral Small 3 24B 24B Apple M5 (32GB)~16.3 GB
- 4 1347 Qwen3 32B 32B Nvidia GeForce RTX 4090 (24GB)~22 GB
- 5 1342 Gemma 3 12B 12B Apple M2 (16GB)~8.9 GB
- 6 1318 Llama 3.3 70B 70B Apple M4 Max (64GB)~45.3 GB
- 7 1303 Gemma 3 4B 4B iPhone 15 Pro~3.8 GB
- 8 1303 Qwen2.5 72B 72B Apple M4 Max (128GB)~50.2 GB
- 9 1289 Gemma 2 27B 27B Apple M5 (32GB)~18.7 GB
- 10 1266 Gemma 2 9B 9B iPhone 17 Pro~7.3 GB
- 11 1256 Phi-4 14B 14B Nvidia GeForce RTX 3060 (12GB)~10.8 GB
- 12 1211 Llama 3.1 8B 8B iPhone 17 Pro~6.4 GB
- 13 1166 Llama 3.2 3B 3B iPhone 15 Pro~3.2 GB
- 14 1149 Mistral 7B 7B iPhone 17 Pro~5.8 GB
- 15 1114 SmolLM2 1.7B 1.7B iPhone 15 Pro~2.2 GB
- 16 1110 Llama 3.2 1B 1B iPhone 15 Pro~1.8 GB
Score is LMArena (human-preference elo from lmarena's blind head-to-head votes.), sourced from lmarena.ai/leaderboard. The "runs on" column is the lightest tracked device that runs the model at Q4_K_M; tap a model for the full hardware list. Coverage is partial: only models we verified against the canonical board appear.
Other boards
Or pick by what you need with use-case picks, or browse the full model list.
FAQ
What is the best local model for general chat?
Qwen3 30B-A3B tops this board at 1383 on LMArena. It needs ~20.7 GB at Q4_K_M, so the lightest hardware that runs it is Apple M5 (32GB).
How is this chat (elo) leaderboard scored?
LMArena ranks models by Elo from blind, head-to-head human votes. It captures overall chat quality and instruction-following as people actually perceive it, which is why it is the most-cited general leaderboard. We only list a model after verifying its score against the canonical leaderboard for that exact model, so coverage is partial by design. Source: lmarena.ai/leaderboard.
Can I run these models locally?
Every row shows the lightest tracked device that runs the model at Q4_K_M and links to the full hardware list, so you see the ranking and your hardware fit together.
Benchmark scores are third-party, sourced from lmarena.ai/leaderboard, and reflect the full-precision model, not a specific quant. Memory and fit are computed at Q4_K_M, validated 2026-08-03. See methodology.