Skip to content

Selected comparison

Local vs cloud LLM comparison

These 6 local models need an estimated 13.2–62.4 GB at Q4 and 4K context. Compare them with 4 hosted options from Claude, Astra and Gemini. Choose your hardware, then estimate monthly API token costs alongside Arena's preference ratings from 2026-09-13.

Showing 10 selected models. Local fit uses Apple M4 (24GB) at 4K context.

Usage estimate

Estimate listed API spend

Enter millions of tokens for one month. Only hosted rows with a listed standard rate receive an estimate.

Estimates use standard-tier uncached input and billed output rates. They exclude cache reads and writes, tools, batch or priority discounts, long-request surcharges and taxes. Known expired rates show Unavailable. Follow each price link for vendor notes.

Preference and fit

Selected models, highest rating first

Open source snapshot

Scroll the table sideways for local memory fit and API costs.

Selected local and hosted models with Arena preference ratings, local memory fit and listed API costs.
ModelArena preference ratingLocal fit and Q4 memoryDated API priceMonthly API estimate
Gemini 3.8 Flash (High)gemini-3.8-flash-high · Hosted 1493.01484.4–1501.6 · 5,076 votes Hosted only In $0.75 / MOut $3.75 / MVerified 2026-09-12Through 2026-12-31 $1.50
Claude Opus 5 (High)claude-opus-5-high · Hosted 1492.91488.7–1497.1 · 42,617 votes Hosted only In $5.00 / MOut $25.00 / MVerified 2026-09-12 $10.00
GPT-6 Astra (Max)gpt-6-astra-max · Hosted 1479.81468.1–1491.4 · 2,693 votes Hosted only In $10.00 / MOut $50.00 / MVerified 2026-09-12 $20.00
Claude Sonnet 5 (High)claude-sonnet-5-high · Hosted 1461.21456.6–1465.8 · 35,301 votes Hosted only In $2.00 / MOut $10.00 / MVerified 2026-09-12 $4.00
Gemma 4 31Bgemma-4-31b · Local 1451.11443.5–1458.7 · 5,894 votes No fit~20.5 GB4K setup details Not listed Not listed
Gemma 4 26B-A4Bgemma-4-26b-a4b · Local 1437.91430.3–1445.5 · 5,804 votes No fit~19 GB4K setup details Not listed Not listed
gpt-oss 120Bgpt-oss-120b · Local 1352.31347.9–1356.7 · 29,955 votes No fit~62.4 GB4K setup details Not listed Not listed
Qwen3 32Bqwen3-32b · Local 1346.91337.5–1356.4 · 3,926 votes No fit~22 GB4K setup details Not listed Not listed
Qwen3 30B-A3Bqwen3-30b-a3b · Local 1326.91322.1–1331.6 · 26,089 votes No fit~20.7 GB4K setup details Not listed Not listed
gpt-oss 20Bgpt-oss-20b · Local 1317.21310.8–1323.6 · 10,393 votes Fits~13.2 GB4K setup details Not listed Not listed

The bounds show the source confidence interval. Rating intervals can overlap, so do not assume a reliable ordering from nearby rows. Names retain the source's evaluated reasoning modes. Arena's settings can differ from a local Q4 installation.

How to read this

One source snapshot, two deployment paths

The preference ratings come from the linked Arena snapshot on 2026-09-13. Local memory uses this catalog's Q4 measurements and selected context. Monthly API usage is separate from that per-request context setting. Billed output includes reasoning tokens, which can exceed the visible answer. Memory fit does not guarantee runtime support or usable speed. Vendor API prices have their own verification dates and links per row. These are separate facts, shown together for planning, not evidence of parity.

For a broader local benchmark view, visit the local leaderboard. To size any local model for a machine, use the memory calculator. Can you run Claude locally? explains the open-weight boundary.

Methodology and dates

Arena: Bradley-Terry rating, adapted from the dataset snapshot dated 2026-09-13; rounded to one decimal here.

Memory: Q4 local-catalog estimate recomputed for the selected device and context. Memory data updated 2026-09-14.

Prices: USD per million tokens, separate from subscription plans. Vendor rates below cover the base API model; the Arena label retains the evaluated reasoning mode.

Gemini 3.8 Flash (High) (gemini-3.8-flash), verified 2026-09-12: Standard paid Gemini Developer API rates through 2026-12-31. Published rates from 2027-01-01 are $1.50 input and $7.50 output per million. Output includes thinking tokens. Cache storage, tools and other tiers are excluded.

Claude Opus 5 (High) (claude-opus-5), verified 2026-09-12: Standard global Claude API, including the 1M context window. US-only inference costs 1.1x. Fast mode, cache creation and tools are excluded.

GPT-6 Astra (Max) (gpt-6-astra), verified 2026-09-12: Standard text rates for requests with at most 272K input tokens. Above that threshold the whole request has 2x input and 1.5x output rates. Cache creation and tools are excluded.

Claude Sonnet 5 (High) (claude-sonnet-5), verified 2026-09-12: Standard global Claude API, including the 1M context window. US-only inference costs 1.1x. Cache creation and tools are excluded.

Adapted from Arena leaderboard dataset, CC BY 4.0.

FAQ

Frequently asked questions

Can I run a cloud model locally?

Usually not. Hosted rows represent a provider API run, while local rows link to an open-weight model with a Q4 memory estimate. The hardware selector only evaluates the local catalog models in this selected comparison.

Does an Arena preference rating measure local performance?

No. The rating comes from the listed Arena leaderboard snapshot and retains its evaluated mode. A hosted Arena run is not a local quantization benchmark, so use the rating for preference context and the memory verdict for local hardware planning.

Are the API costs a bill quote?

No. The estimate multiplies your entered input and output millions of tokens by the linked vendor rates in this snapshot. It excludes subscriptions, tools, regional pricing, taxes and any rate that is not listed.

Source retrieved 2026-09-14. Download the comparison data. Selection is editorial and does not claim a global rank. Ratings are rounded for display; source and license links above carry the full dataset context.