Leaderboard · Aider polyglot
Local LLM coding leaderboard
The Aider polyglot benchmark runs models through 225 Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, scoring the share solved in Aider's test setup. This board ranks the tracked models with recorded scores; it does not measure every coding workflow.
Benchmark verification dates are not yet recorded. The model-size refresh does not recheck scores; consult the source leaderboard for current results. Hardware fit below is estimated at Q4_K_M and does not establish the same benchmark performance.
- 1 71.4% DeepSeek-R1-0528 671B MoE big rig~384.1 GB
- 2 59.6% Qwen3 235B A22B 235B MoE Apple M3 Ultra (256GB)~136.9 GB
- 3 59.1% KI Kimi K2 Instruct 1000B MoE big rig~586.6 GB
- 4 56.9% DeepSeek R1 671B MoE big rig~383.7 GB
- 5 55.1% DeepSeek V3 671B MoE big rig~383.7 GB
- 6 41.8% gpt-oss 120B 117B MoE Apple M4 Max (128GB)~62.4 GB
- 7 40% Qwen3 32B 32B Nvidia GeForce RTX 4090 (24GB)~22 GB
- 8 15.6% Llama 4 Maverick 400B MoE big rig~231.7 GB
- 9 12% CR Command A 111B Apple M4 Max (128GB)~70.4 GB
- 10 4.9% Gemma 3 27B 27B Apple M5 (32GB)~18.6 GB
Score is Aider polyglot (percent of 225 exercism exercises solved across six languages.), sourced from aider.chat/docs/leaderboards. The "runs on" column is the lightest tracked device that runs the model at Q4_K_M; tap a model for the full hardware list. Coverage is partial: only models we verified against the canonical board appear.
Other boards
Or pick by what you need with use-case picks, or browse the full model list.
FAQ
Which tracked model has the highest recorded coding score?
DeepSeek-R1-0528 leads our partial catalog at 71.4% on Aider polyglot. This does not establish the current overall leader. It needs ~384.1 GB at Q4_K_M, more than any single tracked device, so it wants a high-memory rig. The hardware figure is an estimate, not a measured benchmark run on that device.
How is this coding leaderboard scored?
The Aider polyglot benchmark runs models through 225 Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, scoring the share solved in Aider's test setup. This board ranks the tracked models with recorded scores; it does not measure every coding workflow. We only list a model after verifying its score against the canonical leaderboard for that exact model, so coverage is partial by design. Source: aider.chat/docs/leaderboards.
Can I run these models locally?
Every row shows the lightest tracked device that runs the model at Q4_K_M and links to the full hardware list, so you see the ranking and your hardware fit together.
Benchmark scores come from aider.chat/docs/leaderboards and describe its evaluation setup. Memory-data refresh: 2026-09-14; this is not a benchmark verification date. See methodology.