Runs in the browser · Text generation
Run Llama-3.2-1B-Instruct in your browser
Llama-3.2-1B-Instruct downloads 1.01 GB over WebGPU (q4f16), or 1.15 GB over WebAssembly (uint8), to generate text entirely in the tab with Transformers.js. No install, no server.
Text generation · onnx-community/Llama-3.2-1B-Instruct
- Parameters
- 1.24B
- Pipeline task
- text-generation
Reading
Every build exceeds 1 GB; q4f16 (about 1.04 GB) is the lightest WebGPU option, and even that is a meaningful download and load-time hit for a browser page.
Will it run in your browser?
Checking for WebGPU support…
All measured variants
| Variant | Size |
|---|---|
| q4f16 WebGPU pick | 1.01 GB |
| uint8 WASM pick | 1.15 GB |
| bnb4 | 1.49 GB |
| q4 | 1.58 GB |
| fp32 | 1.94 GB |
| fp16 | 1.95 GB |
| q8 | 2.30 GB |
Sizes measured from the HuggingFace API file tree for onnx-community/Llama-3.2-1B-Instruct, not estimated.
Use it with Transformers.js
import { pipeline } from "@huggingface/transformers";
const pipe = await pipeline("text-generation", "onnx-community/Llama-3.2-1B-Instruct", {
device: "webgpu", // usually falls back to "wasm"; wrap in try/catch for production
dtype: "q4f16", // use "uint8" for the WASM build
});
const result = await pipe(/* your input */);
Requires npm install @huggingface/transformers (or the CDN build).
Sources
Same job, different size
FAQ
WebGPU or WASM for Llama-3.2-1B-Instruct, in practice?
Chrome, Edge, and Safari 26+ run the WebGPU build (q4f16) on the GPU. Firefox has WebGPU on Windows and Apple Silicon Macs but not everywhere yet. Any browser without a WebGPU adapter falls back automatically to the WebAssembly build (uint8) on CPU: same model, slower to load, slower to run.
What does q4f16 mean for Llama-3.2-1B-Instruct?
The WebGPU build here uses q4f16: 4-bit weights with some layers kept at 16-bit for stability: WebGPU's usual smallest clean build. The WASM fallback uses uint8: 8-bit integers: about a quarter the size of fp32, with a smaller quality trade than 4-bit.
Sizes measured 2026-08-01 from the HuggingFace API. Last validated 2026-08-03. See all browser models or the methodology.