Skip to content

By use case · Vision

Best local vision (multimodal) models

These models read images as input, so you can ask about a screenshot or a photo, not only text. 6 open vision (multimodal) models are ranked by quality below. The strongest is Qwen3-VL 30B-A3B (~20.4 GB at Q4_K_M); the lightest is Qwen3-VL 4B (~4.4 GB).

Models that accept images as input alongside text.

Tagged by each model's stated purpose. The memory figure is what it needs at Q4_K_M, and the device beside it is the lightest tracked machine that fits it. "Runs on N devices" counts the 43 tracked devices that fit it at Q4_K_M.

FAQ

What is the best local vision (multimodal) model?

Qwen3-VL 30B-A3B leads on quality among the open vision (multimodal) models tracked here. It needs ~20.4 GB at Q4_K_M, so the lightest hardware that runs it is Apple M5 (32GB). Pick by what fits your memory using the list above.

What is the smallest vision (multimodal) model that runs on a laptop?

Qwen3-VL 4B is the lightest, at ~4.4 GB at Q4_K_M, so any device with 8 GB or more can load it.

How were these vision (multimodal) models chosen?

They are open-weight models whose own design targets vision (multimodal) (by name, family or model card). Memory figures are computed at Q4_K_M and sourced; see the methodology page.

Memory is computed at Q4_K_M; catalog updated 2026-10-05. See methodology.