Selected comparison
Local vs cloud LLM comparison
These 6 local models need an estimated 13.2–62.4 GB at Q4 and 4K context. Compare them with 4 hosted options from Claude, Astra and Gemini. Choose your hardware, then estimate monthly API token costs alongside Arena's preference ratings from 2026-09-13.
Showing 10 selected models. Local fit uses Apple M4 (24GB) at 4K context.
Usage estimate
Estimate listed API spend
Enter millions of tokens for one month. Only hosted rows with a listed standard rate receive an estimate.
Estimates use standard-tier uncached input and billed output rates. They exclude cache reads and writes, tools, batch or priority discounts, long-request surcharges and taxes. Known expired rates show Unavailable. Follow each price link for vendor notes.
Preference and fit
Selected models, highest rating first
Scroll the table sideways for local memory fit and API costs.
| Model | Arena preference rating | Local fit and Q4 memory | Dated API price | Monthly API estimate |
|---|---|---|---|---|
| Gemini 3.8 Flash (High)gemini-3.8-flash-high · Hosted | 1493.01484.4–1501.6 · 5,076 votes | Hosted only | In $0.75 / MOut $3.75 / MVerified 2026-09-12Through 2026-12-31 | $1.50 |
| Claude Opus 5 (High)claude-opus-5-high · Hosted | 1492.91488.7–1497.1 · 42,617 votes | Hosted only | In $5.00 / MOut $25.00 / MVerified 2026-09-12 | $10.00 |
| GPT-6 Astra (Max)gpt-6-astra-max · Hosted | 1479.81468.1–1491.4 · 2,693 votes | Hosted only | In $10.00 / MOut $50.00 / MVerified 2026-09-12 | $20.00 |
| Claude Sonnet 5 (High)claude-sonnet-5-high · Hosted | 1461.21456.6–1465.8 · 35,301 votes | Hosted only | In $2.00 / MOut $10.00 / MVerified 2026-09-12 | $4.00 |
| Gemma 4 31Bgemma-4-31b · Local | 1451.11443.5–1458.7 · 5,894 votes | No fit~20.5 GB4K setup details | Not listed | Not listed |
| Gemma 4 26B-A4Bgemma-4-26b-a4b · Local | 1437.91430.3–1445.5 · 5,804 votes | No fit~19 GB4K setup details | Not listed | Not listed |
| gpt-oss 120Bgpt-oss-120b · Local | 1352.31347.9–1356.7 · 29,955 votes | No fit~62.4 GB4K setup details | Not listed | Not listed |
| Qwen3 32Bqwen3-32b · Local | 1346.91337.5–1356.4 · 3,926 votes | No fit~22 GB4K setup details | Not listed | Not listed |
| Qwen3 30B-A3Bqwen3-30b-a3b · Local | 1326.91322.1–1331.6 · 26,089 votes | No fit~20.7 GB4K setup details | Not listed | Not listed |
| gpt-oss 20Bgpt-oss-20b · Local | 1317.21310.8–1323.6 · 10,393 votes | Fits~13.2 GB4K setup details | Not listed | Not listed |
The bounds show the source confidence interval. Rating intervals can overlap, so do not assume a reliable ordering from nearby rows. Names retain the source's evaluated reasoning modes. Arena's settings can differ from a local Q4 installation.
How to read this
One source snapshot, two deployment paths
The preference ratings come from the linked Arena snapshot on 2026-09-13. Local memory uses this catalog's Q4 measurements and selected context. Monthly API usage is separate from that per-request context setting. Billed output includes reasoning tokens, which can exceed the visible answer. Memory fit does not guarantee runtime support or usable speed. Vendor API prices have their own verification dates and links per row. These are separate facts, shown together for planning, not evidence of parity.
For a broader local benchmark view, visit the local leaderboard. To size any local model for a machine, use the memory calculator. Can you run Claude locally? explains the open-weight boundary.
Methodology and dates
Arena: Bradley-Terry rating, adapted from the dataset snapshot dated 2026-09-13; rounded to one decimal here.
Memory: Q4 local-catalog estimate recomputed for the selected device and context. Memory data updated 2026-09-14.
Prices: USD per million tokens, separate from subscription plans. Vendor rates below cover the base API model; the Arena label retains the evaluated reasoning mode.
Gemini 3.8 Flash (High) (gemini-3.8-flash), verified 2026-09-12: Standard paid Gemini Developer API rates through 2026-12-31. Published rates from 2027-01-01 are $1.50 input and $7.50 output per million. Output includes thinking tokens. Cache storage, tools and other tiers are excluded.
Claude Opus 5 (High) (claude-opus-5), verified 2026-09-12: Standard global Claude API, including the 1M context window. US-only inference costs 1.1x. Fast mode, cache creation and tools are excluded.
GPT-6 Astra (Max) (gpt-6-astra), verified 2026-09-12: Standard text rates for requests with at most 272K input tokens. Above that threshold the whole request has 2x input and 1.5x output rates. Cache creation and tools are excluded.
Claude Sonnet 5 (High) (claude-sonnet-5), verified 2026-09-12: Standard global Claude API, including the 1M context window. US-only inference costs 1.1x. Cache creation and tools are excluded.
FAQ
Frequently asked questions
Can I run a cloud model locally?
Usually not. Hosted rows represent a provider API run, while local rows link to an open-weight model with a Q4 memory estimate. The hardware selector only evaluates the local catalog models in this selected comparison.
Does an Arena preference rating measure local performance?
No. The rating comes from the listed Arena leaderboard snapshot and retains its evaluated mode. A hosted Arena run is not a local quantization benchmark, so use the rating for preference context and the memory verdict for local hardware planning.
Are the API costs a bill quote?
No. The estimate multiplies your entered input and output millions of tokens by the linked vendor rates in this snapshot. It excludes subscriptions, tools, regional pricing, taxes and any rate that is not listed.
Source retrieved 2026-09-14. Download the comparison data. Selection is editorial and does not claim a global rank. Ratings are rounded for display; source and license links above carry the full dataset context.