Skip to content

Leaderboard

The local LLM leaderboard

Open-weight models ranked by sourced benchmarks, coding, tool use, and chat. The difference: every ranked model shows the lightest hardware that runs it at Q4_K_M, so the ranking and your hardware fit live in one place.

Every benchmarked model

31 models, every sourced score Q4_K_M fit
Every benchmarked model with its coding, tool-use and chat scores, and the lightest device that runs it.
Model Code Tools Elo Runs on
KI Kimi K2 Instruct 1000B 59.1% 59.1% 586.6 GB+
DeepSeek R1 671B 56.9% 383.7 GB+
DeepSeek V3 671B 55.1% 383.7 GB+
DeepSeek-R1-0528 671B 71.4% 384.1 GB+
Llama 4 Maverick 400B 15.6% 37.3% 231.7 GB+
Qwen3 235B A22B 235B 59.6% Apple M3 Ultra (256GB)
gpt-oss 120B 117B 41.8% Apple M4 Max (128GB)
CR Command A 111B 12% 46.5% Apple M4 Max (128GB)
Llama 4 Scout 109B 28.1% Apple M4 Max (128GB)
Qwen2.5 72B 72B 1303 Apple M4 Max (128GB)
Llama 3.3 70B 70B 31.9% 1318 Apple M4 Max (64GB)
Qwen3 32B 32B 40% 48.7% 1347 Nvidia GeForce RTX 4090 (24GB)
Qwen3 30B-A3B 30.5B 1383 Apple M5 (32GB)
Gemma 2 27B 27B 1289 Apple M5 (32GB)
Gemma 3 27B 27B 4.9% 29.5% 1366 Apple M5 (32GB)
Mistral Small 3 24B 24B 1357 Apple M5 (32GB)
Phi-4 14B 14B 28.8% 1256 Nvidia GeForce RTX 3060 (12GB)
Qwen3 14B 14B 41% Nvidia GeForce RTX 3060 (12GB)
Gemma 3 12B 12B 30.4% 1342 Apple M2 (16GB)
FN Falcon3 10B 10B 27% iPhone 17 Pro
Gemma 2 9B 9B 1266 iPhone 17 Pro
Llama 3.1 8B 8B 25.8% 1211 iPhone 17 Pro
Qwen3 8B 8B 42.6% iPhone 17 Pro
Mistral 7B 7B 1149 iPhone 17 Pro
Gemma 3 4B 4B 19.6% 1303 iPhone 15 Pro
Llama 3.2 3B 3B 22% 1166 iPhone 15 Pro
SmolLM2 1.7B 1.7B 1114 iPhone 15 Pro
Qwen3 1.7B 1.7B 28.4% iPhone 15 Pro
Llama 3.2 1B 1B 10.8% 1110 iPhone 15 Pro
Gemma 3 1B 1B 7.2% iPhone 15 Pro
Qwen3 0.6B 600M 23.9% iPhone 15 Pro

A model appears once it has at least one verified score; a dot means no checked score on that board yet (coverage is partial by design, never estimated). Scores are full-precision; memory and the "runs on" device are computed at Q4_K_M. Sorted by size, biggest first. See each board above for the ranked view.

FAQ

What is the best local LLM right now?

It depends what for. For coding, DeepSeek-R1-0528 leads the Aider polyglot board at 71.4%. For tool use and agents, Kimi K2 Instruct tops BFCL at 59.1%. For general chat, Qwen3 30B-A3B has the highest LMArena Elo at 1383.

Why is this leaderboard different?

Every other leaderboard stops at the score. This one ties each ranked model to whether you can actually run it: each row shows the lightest tracked device that fits the model at Q4_K_M, and links to the full hardware breakdown.

Where do the scores come from?

Each board is a sourced, third-party benchmark: Aider polyglot for coding, the Berkeley Function-Calling Leaderboard for tool use, and LMArena for chat Elo. We only list a model after verifying its score against the canonical board for that exact model, so coverage is partial by design and never estimated.

Benchmark scores are third-party and sourced (Aider, BFCL, LMArena); we list a model only after verifying its score against the canonical board. Memory and fit are computed, validated 2026-08-03. See methodology.