Runs in the browser · Text generation
Run SmolLM2-135M-Instruct in your browser
SmolLM2-135M-Instruct downloads 112.2 MB over WebGPU (q4f16), or 130.8 MB over WebAssembly (uint8), to generate text entirely in the tab with Transformers.js. No install, no server.
Text generation · HuggingFaceTB/SmolLM2-135M-Instruct
- Parameters
- 134.5M
- Pipeline task
- text-generation
Reading
The smallest text generation download in the catalog (112.2 MB).
The quant ladder spans 4.6x: 112.2 MB (q4f16) to 515.3 MB (fp32).
Run it in your browser
Checking what this browser can run…
All measured variants
| Variant | Size |
|---|---|
| q4f16 WebGPU pick | 112.2 MB |
| uint8 WASM pick | 130.8 MB |
| bnb4 | 167.3 MB |
| q4 | 173.6 MB |
| fp16 | 257.7 MB |
| q8 | 261.6 MB |
| fp32 | 515.3 MB |
Sizes measured from the HuggingFace API file tree for HuggingFaceTB/SmolLM2-135M-Instruct, not estimated.
Use it with Transformers.js
import { pipeline } from "@huggingface/transformers";
const pipe = await pipeline("text-generation", "HuggingFaceTB/SmolLM2-135M-Instruct", {
device: "webgpu", // usually falls back to "wasm"; wrap in try/catch for production
dtype: "q4f16", // use "uint8" for the WASM build
});
const result = await pipe(/* your input */);
Requires npm install @huggingface/transformers (or the CDN build).
Sources
Same job, different size
FAQ
WebGPU or WASM for SmolLM2-135M-Instruct, in practice?
Chrome, Edge, and Safari 26+ run the WebGPU build (q4f16) on the GPU. Firefox has WebGPU on Windows and Apple Silicon Macs but not everywhere yet. Any browser without a WebGPU adapter falls back automatically to the WebAssembly build (uint8) on CPU: same model, slower to load, slower to run.
What does q4f16 mean for SmolLM2-135M-Instruct?
The WebGPU build here uses q4f16: 4-bit weights with some layers kept at 16-bit for stability: WebGPU's usual smallest clean build. The WASM fallback uses uint8: 8-bit integers: about a quarter the size of fp32, with a smaller quality trade than 4-bit.
Sizes measured 2026-08-01 from the HuggingFace API. Last validated 2026-08-03. See all browser models or the methodology.