Skip to content

Leaderboard · Aider polyglot

Best local LLMs for coding

The Aider polyglot benchmark runs each model through 225 hard Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, scoring the share it solves. It is a real coding signal, not a chat-preference vote, so it ranks the models you would actually pair-program with.

Score is Aider polyglot (percent of 225 exercism exercises solved across six languages.), sourced from aider.chat/docs/leaderboards. The "runs on" column is the lightest tracked device that runs the model at Q4_K_M; tap a model for the full hardware list. Coverage is partial: only models we verified against the canonical board appear.

Other boards

Or pick by what you need with use-case picks, or browse the full model list.

FAQ

What is the best local model for coding?

DeepSeek-R1-0528 tops this board at 71.4% on Aider polyglot. It needs ~384.1 GB at Q4_K_M, more than any single tracked device, so it wants a high-memory rig.

How is this coding leaderboard scored?

The Aider polyglot benchmark runs each model through 225 hard Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, scoring the share it solves. It is a real coding signal, not a chat-preference vote, so it ranks the models you would actually pair-program with. We only list a model after verifying its score against the canonical leaderboard for that exact model, so coverage is partial by design. Source: aider.chat/docs/leaderboards.

Can I run these models locally?

Every row shows the lightest tracked device that runs the model at Q4_K_M and links to the full hardware list, so you see the ranking and your hardware fit together.

Benchmark scores are third-party, sourced from aider.chat/docs/leaderboards, and reflect the full-precision model, not a specific quant. Memory and fit are computed at Q4_K_M, validated 2026-08-03. See methodology.