Skip to content

Leaderboard · LMArena

Local LLM chat leaderboard

LMArena uses blind, head-to-head human votes to measure chat preferences. This board shows the scores recorded for tracked models. Ratings depend on the evaluated model pool and snapshot, so they are not a fixed measure of capability.

Benchmark verification dates are not yet recorded. The model-size refresh does not recheck scores; consult the source leaderboard for current results. Hardware fit below is estimated at Q4_K_M and does not establish the same benchmark performance.

Score is LMArena (human-preference elo from lmarena's blind head-to-head votes.), sourced from lmarena.ai/leaderboard. The "runs on" column is the lightest tracked device that runs the model at Q4_K_M; tap a model for the full hardware list. Coverage is partial: only models we verified against the canonical board appear.

Other boards

Or pick by what you need with use-case picks, or browse the full model list.

FAQ

Which tracked model has the highest recorded chat score?

Qwen3 30B-A3B leads our partial catalog at 1383 on LMArena. This does not establish the current overall leader. It needs ~20.7 GB at Q4_K_M, so the lightest hardware that runs it is Apple M5 (32GB). The hardware figure is an estimate, not a measured benchmark run on that device.

How is this chat (elo) leaderboard scored?

LMArena uses blind, head-to-head human votes to measure chat preferences. This board shows the scores recorded for tracked models. Ratings depend on the evaluated model pool and snapshot, so they are not a fixed measure of capability. We only list a model after verifying its score against the canonical leaderboard for that exact model, so coverage is partial by design. Source: lmarena.ai/leaderboard.

Can I run these models locally?

Every row shows the lightest tracked device that runs the model at Q4_K_M and links to the full hardware list, so you see the ranking and your hardware fit together.

Benchmark scores come from lmarena.ai/leaderboard and describe its evaluation setup. Memory-data refresh: 2026-09-14; this is not a benchmark verification date. See methodology.