Skip to content

Leaderboard

The local LLM leaderboard

Open-weight models ranked by sourced benchmarks, coding, tool use, and chat. The difference: every ranked model shows the lightest hardware that runs it at Q4_K_M, so the ranking and your hardware fit live in one place.

These are recorded third-party results with partial coverage. Benchmark verification dates are not yet tracked; the catalog refresh updates model sizes, not these scores. Open each source board for its latest results.

Compare local models with Astra, Claude and Gemini using a separate, dated Arena snapshot plus local memory fit and API prices. Considering a local coding setup? Read Can you run Claude locally? for the distinction between Claude models, Claude Code and local alternatives.

Every benchmarked model

31 models, every sourced score Q4_K_M fit
Every benchmarked model with its coding, tool-use and chat scores, and the lightest device that runs it.
Model Code Tools Elo Runs on
KI Kimi K2 Instruct 1000B 59.1% 59.1% 586.6 GB+
DeepSeek R1 671B 56.9% 383.7 GB+
DeepSeek V3 671B 55.1% 383.7 GB+
DeepSeek-R1-0528 671B 71.4% 384.1 GB+
Llama 4 Maverick 400B 15.6% 37.3% 231.7 GB+
Qwen3 235B A22B 235B 59.6% Apple M3 Ultra (256GB)
gpt-oss 120B 117B 41.8% Apple M4 Max (128GB)
CR Command A 111B 12% 46.5% Apple M4 Max (128GB)
Llama 4 Scout 109B 28.1% Apple M4 Max (128GB)
Qwen2.5 72B 72B 1303 Apple M4 Max (128GB)
Llama 3.3 70B 70B 31.9% 1318 Apple M4 Max (64GB)
Qwen3 32B 32B 40% 48.7% 1347 Nvidia GeForce RTX 4090 (24GB)
Qwen3 30B-A3B 30.5B 1383 Apple M5 (32GB)
Gemma 2 27B 27B 1289 Apple M5 (32GB)
Gemma 3 27B 27B 4.9% 29.5% 1366 Apple M5 (32GB)
Mistral Small 3 24B 24B 1357 Apple M5 (32GB)
Phi-4 14B 14B 28.8% 1256 Nvidia GeForce RTX 3060 (12GB)
Qwen3 14B 14B 41% Nvidia GeForce RTX 3060 (12GB)
Gemma 3 12B 12B 30.4% 1342 Apple M2 (16GB)
FN Falcon3 10B 10B 27% iPhone 17 Pro
Gemma 2 9B 9B 1266 iPhone 17 Pro
Llama 3.1 8B 8B 25.8% 1211 Nvidia GeForce RTX 3060 Ti (8GB)
Qwen3 8B 8B 42.6% Nvidia GeForce RTX 3060 Ti (8GB)
Mistral 7B 7B 1149 Nvidia GeForce RTX 3060 Ti (8GB)
Gemma 3 4B 4B 19.6% 1303 iPhone 15 Pro
Llama 3.2 3B 3B 22% 1166 iPhone 15 Pro
SmolLM2 1.7B 1.7B 1114 iPhone 15 Pro
Qwen3 1.7B 1.7B 28.4% iPhone 15 Pro
Llama 3.2 1B 1B 10.8% 1110 iPhone 15 Pro
Gemma 3 1B 1B 7.2% iPhone 15 Pro
Qwen3 0.6B 600M 23.9% iPhone 15 Pro

A model appears once it has at least one recorded score; a dot means no checked score on that board yet. Scores describe the source's evaluation setup, not a measured Q4_K_M run on the listed device. Memory and the "runs on" device are estimates at Q4_K_M. Sorted by size, biggest first. See each board above for the ranked view.

FAQ

Which tracked models have the highest recorded scores?

Among the models with scores in our catalog, DeepSeek-R1-0528 leads coding at 71.4% on Aider polyglot. Kimi K2 Instruct has the highest recorded BFCL score at 59.1%. Qwen3 30B-A3B has the highest recorded chat Elo at 1383. Coverage is partial and score verification dates are not yet recorded, so these are not claims about the current overall leaders.

Why is this leaderboard different?

Each ranked model links to a hardware estimate: the lightest tracked device that fits it at Q4_K_M, plus the full memory breakdown. A hardware fit does not mean that a local quantized run will reproduce the benchmark score.

Where do the scores come from?

Each board is a sourced, third-party benchmark: Aider polyglot for coding, the Berkeley Function-Calling Leaderboard for tool use, and LMArena for chat Elo. We only list a model after verifying its score against the canonical board for that exact model, so coverage is partial by design and never estimated.

Sources: Aider polyglot, BFCL and LMArena. Memory-data refresh: 2026-09-14; this is not a benchmark verification date. See methodology.