Skip to content

Leaderboard · LMArena

Best local LLMs by human preference

LMArena ranks models by Elo from blind, head-to-head human votes. It captures overall chat quality and instruction-following as people actually perceive it, which is why it is the most-cited general leaderboard.

Score is LMArena (human-preference elo from lmarena's blind head-to-head votes.), sourced from lmarena.ai/leaderboard. The "runs on" column is the lightest tracked device that runs the model at Q4_K_M; tap a model for the full hardware list. Coverage is partial: only models we verified against the canonical board appear.

Other boards

Or pick by what you need with use-case picks, or browse the full model list.

FAQ

What is the best local model for general chat?

Qwen3 30B-A3B tops this board at 1383 on LMArena. It needs ~20.7 GB at Q4_K_M, so the lightest hardware that runs it is Apple M5 (32GB).

How is this chat (elo) leaderboard scored?

LMArena ranks models by Elo from blind, head-to-head human votes. It captures overall chat quality and instruction-following as people actually perceive it, which is why it is the most-cited general leaderboard. We only list a model after verifying its score against the canonical leaderboard for that exact model, so coverage is partial by design. Source: lmarena.ai/leaderboard.

Can I run these models locally?

Every row shows the lightest tracked device that runs the model at Q4_K_M and links to the full hardware list, so you see the ranking and your hardware fit together.

Benchmark scores are third-party, sourced from lmarena.ai/leaderboard, and reflect the full-precision model, not a specific quant. Memory and fit are computed at Q4_K_M, validated 2026-08-03. See methodology.