Leaderboard
The local LLM leaderboard
Open-weight models ranked by sourced benchmarks, coding, tool use, and chat. The difference: every ranked model shows the lightest hardware that runs it at Q4_K_M, so the ranking and your hardware fit live in one place.
Every benchmarked model
A model appears once it has at least one verified score; a dot means no checked score on that board yet (coverage is partial by design, never estimated). Scores are full-precision; memory and the "runs on" device are computed at Q4_K_M. Sorted by size, biggest first. See each board above for the ranked view.
FAQ
What is the best local LLM right now?
It depends what for. For coding, DeepSeek-R1-0528 leads the Aider polyglot board at 71.4%. For tool use and agents, Kimi K2 Instruct tops BFCL at 59.1%. For general chat, Qwen3 30B-A3B has the highest LMArena Elo at 1383.
Why is this leaderboard different?
Every other leaderboard stops at the score. This one ties each ranked model to whether you can actually run it: each row shows the lightest tracked device that fits the model at Q4_K_M, and links to the full hardware breakdown.
Where do the scores come from?
Each board is a sourced, third-party benchmark: Aider polyglot for coding, the Berkeley Function-Calling Leaderboard for tool use, and LMArena for chat Elo. We only list a model after verifying its score against the canonical board for that exact model, so coverage is partial by design and never estimated.
Benchmark scores are third-party and sourced (Aider, BFCL, LMArena); we list a model only after verifying its score against the canonical board. Memory and fit are computed, validated 2026-08-03. See methodology.