54 models · Transformers.js · WebGPU / WASM
Local AI that runs in the browser
54 models, one tab each: no install, no server, no Ollama. They run with Transformers.js over WebGPU or WebAssembly, and every download size below is measured from the HuggingFace API, not estimated.
Running models outside the browser? See the main GGUF catalog for native runtimes (Ollama, llama.cpp, LM Studio) across desktop and mobile hardware.
- Moonshine Tiny 27.09M params · automatic-speech-recognition The WASM build downloads smaller here: 26.8 MB (uint8) against 53.4 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 53.4 MB WebGPU · q4f16
- Whisper Tiny 37.8M params · automatic-speech-recognition The WASM build downloads smaller here: 39.0 MB (uint8) against 72.6 MB (fp16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 72.6 MB WebGPU · fp16
- Moonshine Base 61.5M params · automatic-speech-recognition The WASM build downloads smaller here: 60.0 MB (uint8) against 97.0 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 97.0 MB WebGPU · q4f16
- Whisper Base 72.6M params · automatic-speech-recognition The WASM build downloads smaller here: 73.3 MB (uint8) against 139.3 MB (fp16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 139.3 MB WebGPU · fp16
- Whisper Small 241.7M params · automatic-speech-recognition The WASM build downloads smaller here: 237.5 MB (uint8) against 462.7 MB (fp16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 462.7 MB WebGPU · fp16
- Whisper Large v3 Turbo 808.88M params · automatic-speech-recognition The second-largest speech to text download in the catalog (537.4 MB). 537.4 MB WebGPU · q4f16
- Voxtral Mini 4B Realtime 4.43B params · automatic-speech-recognition · large download The largest speech to text download in the catalog (2.64 GB). 2.64 GB WebGPU · q4f16
- SmolLM2-135M-Instruct 134.5M params · text-generation The smallest text generation download in the catalog (112.2 MB). 112.2 MB WebGPU · q4f16
- LFM2.5-350M 354.48M params · text-generation The second-smallest text generation download in the catalog (243.3 MB). 243.3 MB WebGPU · q4f16
- SmolLM2-360M-Instruct 361.8M params · text-generation The quant ladder spans 5.3x: 260.1 MB (q4f16) to 1.35 GB (fp32). 260.1 MB WebGPU · q4f16
- Gemma-3-270M-it 268.1M params · text-generation 0.97 MB/M params against a 0.73 MB/M text generation median: dense for its parameter count. 260.3 MB WebGPU · q4f16
- Granite-4.0-350M 352.38M params · text-generation 0.95 MB/M params against a 0.73 MB/M text generation median: dense for its parameter count. 334.2 MB WebGPU · q4f16
- Qwen2.5-0.5B-Instruct 494M params · text-generation 0.93 MB/M params against a 0.73 MB/M text generation median: dense for its parameter count. 460.6 MB WebGPU · q4f16
- Qwen3-0.6B 751.63M params · text-generation 543.4 MB WebGPU · q4f16
- Gemma-3-1B-it 999.89M params · text-generation The quant ladder spans 3.3x: 728.2 MB (q4f16) to 2.35 GB (q8). 728.2 MB WebGPU · q4f16
- Llama-3.2-1B-Instruct 1.24B params · text-generation 1.01 GB WebGPU · q4f16
- Bonsai-1.7B 1.72B params · text-generation · large download 1.04 GB WebGPU · q4
- Qwen2.5-1.5B-Instruct 1.54B params · text-generation · large download The quant ladder spans 5.1x: 1.14 GB (q4f16) to 5.78 GB (fp32). 1.14 GB WebGPU · q4f16
- DeepSeek-R1-Distill-Qwen-1.5B 1.78B params · text-generation · large download The second-largest text generation download in the catalog (1.28 GB). 1.28 GB WebGPU · q4f16
- Phi-3.5-mini-instruct 3.82B params · text-generation · large download The largest text generation download in the catalog (2.16 GB). 2.16 GB WebGPU · q4f16
- all-MiniLM-L6-v2 22.7M params · feature-extraction The WASM build downloads smaller here: 21.8 MB (uint8) against 28.6 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 28.6 MB WebGPU · q4f16
- bge-small-en-v1.5 33.4M params · feature-extraction The WASM build downloads smaller here: 32.2 MB (uint8) against 34.5 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 34.5 MB WebGPU · q4f16
- EmbeddingGemma-300M 302.9M params · feature-extraction 0.55 MB/M params against a 1.03 MB/M embeddings median: light for its parameter count. 168.0 MB WebGPU · q4f16
- voyage-4-nano 346.45M params · feature-extraction 0.58 MB/M params against a 1.03 MB/M embeddings median: light for its parameter count. 201.3 MB WebGPU · q4f16
- GTE Multilingual Base 305M params · feature-extraction The WASM build downloads smaller here: 324.6 MB (uint8) against 443.7 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 443.7 MB WebGPU · q4f16
- Qwen3-Embedding-0.6B 595.78M params · feature-extraction The second-largest embeddings download in the catalog (541.2 MB). 541.2 MB WebGPU · q4f16
- BGE-M3 569M params · feature-extraction The WASM build downloads smaller here: 542.1 MB (uint8) against 667.5 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 667.5 MB WebGPU · q4f16
- MODNet 6.5M params · background-removal The WASM build downloads smaller here: 6.3 MB (uint8) against 11.3 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 11.3 MB WebGPU · q4f16
- DINOv3 ViT-S/16 21.6M params · image-feature-extraction The second-smallest vision download in the catalog (14.1 MB). 14.1 MB WebGPU · q4
- Depth Anything V2 Small 24.8M params · depth-estimation 0.74 MB/M params against a 1.73 MB/M vision median: light for its parameter count. 18.2 MB WebGPU · q4f16
- SlimSAM-77 9.7M params · custom setup The WASM build downloads smaller here: 13.1 MB (q8) against 19.8 MB (fp16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 19.8 MB WebGPU · fp16
- SAM 2.1 Hiera Tiny 38.96M params · custom setup 0.82 MB/M params against a 1.73 MB/M vision median: light for its parameter count. 32.1 MB WebGPU · q4f16
- ORMBG background-removal The WASM build downloads smaller here: 42.3 MB (uint8) against 84.0 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 84.0 MB WebGPU · q4f16
- RMBG-1.4 44.1M params · custom setup The WASM build downloads smaller here: 42.3 MB (q8) against 84.1 MB (fp16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 84.1 MB WebGPU · fp16
- BiRefNet Lite 44.36M params · background-removal 2.46 MB/M params against a 1.73 MB/M vision median: dense for its parameter count. 109.2 MB WebGPU · fp16
- Grounding DINO Tiny 172.28M params · zero-shot-object-detection 0.84 MB/M params against a 1.73 MB/M vision median: light for its parameter count. 144.1 MB WebGPU · q4f16
- BEN2 94.63M params · background-removal 2.21 MB/M params against a 1.73 MB/M vision median: dense for its parameter count. 209.0 MB WebGPU · fp16
- CLIP ViT-B/32 151.3M params · zero-shot-image-classification The second-largest vision download in the catalog (240.0 MB). 240.0 MB WebGPU · q4f16
- SigLIP2 Base Patch16-224 375.19M params · zero-shot-image-classification The WASM build downloads smaller here: 721.0 MB (uint8) against 948.7 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 948.7 MB WebGPU · q4f16
- SmolVLM-256M-Instruct 256.5M params · custom setup The smallest vision + language download in the catalog (180.1 MB). 180.1 MB WebGPU · q4f16
- Florence-2-base-ft 231.6M params · custom setup The second-smallest vision + language download in the catalog (213.1 MB). 213.1 MB WebGPU · q4f16
- Granite-Docling-258M 257.52M params · custom setup 0.98 MB/M params against a 0.72 MB/M vision + language median: dense for its parameter count. 252.1 MB WebGPU · q4f16
- Florence-2-large-ft 770.43M params · custom setup The second-largest vision + language download in the catalog (555.3 MB). 555.3 MB WebGPU · q4f16
- Qwen3.5-0.8B 873.44M params · custom setup The largest vision + language download in the catalog (616.9 MB). 616.9 MB WebGPU · q4f16
- Janus-Pro-1B 2.08B params · custom setup · large download 0.89 MB/M params against a 0.63 MB/M multimodal (any-to-any) median: dense for its parameter count. 1.81 GB WebGPU · q4f16
- Gemma-4-E2B-it 5.12B params · custom setup · large download The WASM build downloads smaller here: 2.89 GB (q8) against 3.15 GB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 3.15 GB WebGPU · q4f16
- Gemma-4-E4B-it 8.00B params · custom setup · large download The WASM build downloads smaller here: 3.18 GB (q8) against 4.07 GB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 4.07 GB WebGPU · q4f16
- GLiNER Small v2.1 166M params · token-classification The WASM build downloads smaller here: 174.9 MB (uint8) against 233.9 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 233.9 MB WebGPU · q4f16
- Punctuate All 278.89M params · token-classification The WASM build downloads smaller here: 265.7 MB (uint8) against 413.1 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 413.1 MB WebGPU · q4f16
- Piiranha v1 (PII Detection) 278.23M params · token-classification The WASM build downloads smaller here: 302.5 MB (uint8) against 432.1 MB (q4f16) for WebGPU. The CPU-fallback quant compresses tighter than the GPU pick for this model. 432.1 MB WebGPU · q4f16
FAQ
What does 'runs in the browser' actually mean?
The model weights (converted to ONNX) download once to the browser cache, then inference runs locally in the tab via Transformers.js, using WebGPU if the browser supports it or WebAssembly as a fallback. Nothing is sent to a server after the download.
WebGPU or WASM: which should I use?
WebGPU is faster and usually the smaller download (the q4f16 quant), but it needs browser support. WebGPU now ships by default in current Chrome, Edge, Safari (26+), and Firefox on Windows and Apple Silicon Macs; older browser versions and Firefox on Linux may still lack it. WASM works everywhere but is slower and often needs a larger int8 build. Each model page below has a live check for your current browser.
Are these the same models as the main GGUF catalog?
No. The main catalog tracks GGUF models for native runtimes (Ollama, llama.cpp, LM Studio) on a device you install software on. This page tracks ONNX exports built specifically for Transformers.js in a browser tab, a different runtime with its own size measurements.
Sizes measured 2026-08-01 from the HuggingFace API file tree for each repo. Last validated 2026-08-03. See methodology.