Leaderboard
The local LLM leaderboard
Open-weight models ranked by sourced benchmarks, coding, tool use, and chat. The difference: every ranked model shows the lightest hardware that runs it at Q4_K_M, so the ranking and your hardware fit live in one place.
These are recorded third-party results with partial coverage. Benchmark verification dates are not yet tracked; the catalog refresh updates model sizes, not these scores. Open each source board for its latest results.
Compare local models with Astra, Claude and Gemini using a separate, dated Arena snapshot plus local memory fit and API prices. Considering a local coding setup? Read Can you run Claude locally? for the distinction between Claude models, Claude Code and local alternatives.
Every benchmarked model
A model appears once it has at least one recorded score; a dot means no checked score on that board yet. Scores describe the source's evaluation setup, not a measured Q4_K_M run on the listed device. Memory and the "runs on" device are estimates at Q4_K_M. Sorted by size, biggest first. See each board above for the ranked view.
FAQ
Which tracked models have the highest recorded scores?
Among the models with scores in our catalog, DeepSeek-R1-0528 leads coding at 71.4% on Aider polyglot. Kimi K2 Instruct has the highest recorded BFCL score at 59.1%. Qwen3 30B-A3B has the highest recorded chat Elo at 1383. Coverage is partial and score verification dates are not yet recorded, so these are not claims about the current overall leaders.
Why is this leaderboard different?
Each ranked model links to a hardware estimate: the lightest tracked device that fits it at Q4_K_M, plus the full memory breakdown. A hardware fit does not mean that a local quantized run will reproduce the benchmark score.
Where do the scores come from?
Each board is a sourced, third-party benchmark: Aider polyglot for coding, the Berkeley Function-Calling Leaderboard for tool use, and LMArena for chat Elo. We only list a model after verifying its score against the canonical board for that exact model, so coverage is partial by design and never estimated.
Sources: Aider polyglot, BFCL and LMArena. Memory-data refresh: 2026-09-14; this is not a benchmark verification date. See methodology.