Runs in the browser · Text generation
Run Gemma-3-1B-it in your browser
Gemma-3-1B-it downloads 728.2 MB over WebGPU (q4f16), or 955.1 MB over WebAssembly (uint8), to generate text entirely in the tab with Transformers.js. No install, no server.
Text generation · onnx-community/gemma-3-1b-it-ONNX
- Parameters
- 999.89M
- Pipeline task
- text-generation
Reading
The quant ladder spans 3.3x: 728.2 MB (q4f16) to 2.35 GB (q8).
Will it run in your browser?
Checking for WebGPU support…
All measured variants
| Variant | Size |
|---|---|
| q4f16 WebGPU pick | 728.2 MB |
| q4 | 819.6 MB |
| uint8 WASM pick | 955.1 MB |
| bnb4 | 1.49 GB |
| fp16 | 1.89 GB |
| fp32 | 1.93 GB |
| q8 | 2.35 GB |
Sizes measured from the HuggingFace API file tree for onnx-community/gemma-3-1b-it-ONNX, not estimated.
Use it with Transformers.js
import { pipeline } from "@huggingface/transformers";
const pipe = await pipeline("text-generation", "onnx-community/gemma-3-1b-it-ONNX", {
device: "webgpu", // usually falls back to "wasm"; wrap in try/catch for production
dtype: "q4f16", // use "uint8" for the WASM build
});
const result = await pipe(/* your input */);
Requires npm install @huggingface/transformers (or the CDN build).
Sources
Same job, different size
FAQ
WebGPU or WASM for Gemma-3-1B-it, in practice?
Chrome, Edge, and Safari 26+ run the WebGPU build (q4f16) on the GPU. Firefox has WebGPU on Windows and Apple Silicon Macs but not everywhere yet. Any browser without a WebGPU adapter falls back automatically to the WebAssembly build (uint8) on CPU: same model, slower to load, slower to run.
What does q4f16 mean for Gemma-3-1B-it?
The WebGPU build here uses q4f16: 4-bit weights with some layers kept at 16-bit for stability: WebGPU's usual smallest clean build. The WASM fallback uses uint8: 8-bit integers: about a quarter the size of fp32, with a smaller quality trade than 4-bit.
Sizes measured 2026-08-01 from the HuggingFace API. Last validated 2026-08-03. See all browser models or the methodology.