Skip to content

2 models · Text to speech · Transformers.js

Text to speech in the browser

2 models handle text to speech in the browser today, from 147.4 MB (Kokoro-82M) to 250.7 MB (Supertonic TTS 2), a 199.1 MB download at the midpoint. Every one runs with Transformers.js over WebGPU or WebAssembly: no server, no install, no account.

Models
2
Smallest
147.4 MB
Median
199.1 MB
Largest
250.7 MB

Every text to speech model, ranked by WebGPU size

Model WebGPU · WASM
Model WebGPU download WASM download
Kokoro-82M 82M params 147.4 MB q4f16 88.1 MB q8
Supertonic TTS 2 250.7 MB fp32 250.7 MB fp32

Sizes measured from the HuggingFace API file tree for each model's repo, not estimated. Sorted smallest to largest by the WebGPU headline pick.

Size ladder

The WebGPU download spans 1.7x here: 147.4 MB (Kokoro-82M) to 250.7 MB (Supertonic TTS 2).

1 of 2 models (Kokoro-82M) download a smaller build over WebAssembly than WebGPU: the CPU-fallback quant compresses tighter than the GPU pick there.

None of the 2 models here fit under 100 MB over WebGPU.

A bytes-per-parameter comparison needs at least 3 models with a verified parameter count and a real spread between them; this group does not clear that bar yet.

Notable text to speech models

Other tasks

Or see every browser model grouped by task, or the main GGUF catalog for native runtimes outside the browser.

FAQ

Do all text to speech models here use the same WebGPU quant?

No. 1 use q4f16, 1 use fp32. q4f16 is the most common pick.

What would it cost to download every text to speech model here?

398.1 MB total over WebGPU across all 2 models, if you tried every one back to back. Most projects only need the single model that fits the job, not the whole set.

Sizes measured 2026-08-01 from the HuggingFace API. Last validated 2026-09-14. See all browser models or the methodology.