toolready. Text to Speech

Text to Speech

Speak text with built-in voices, or a 15M local model.

What this does

Reads text aloud, with a choice of two engines behind the same text box. The default one borrows the voices already installed on your operating system. The other downloads a small neural model once and then synthesises speech locally, which is the one to reach for when you want a file rather than a playback.

Which engine should I pick?

Browser voices uses the Web Speech API — instant, no download, and it gives you Rate and Pitch from 0.5 to 2 plus a Volume slider, with Speak, Pause, Resume and Stop. The catch is that the voice list is whatever your OS provides: rich on macOS and iOS, adequate on Windows, sometimes empty on Linux. The picker sorts voices matching your browser language first and hides the macOS novelty set (Bubbles, Zarvox, Trinoids and friends) so the list stays usable.

KittenTTS is a 15M-parameter model run through ONNX Runtime Web in a worker thread. Voices sound consistent on every platform, and unlike the Web Speech path it produces an audio buffer you can keep.

How do I run the local model?

  1. Switch to the KittenTTS (15M, local) tab and press Download & load model.
  2. Wait for the progress bar. About 60 MB of weights and voice embeddings come from Hugging Face, and the ONNX WASM runtime from jsDelivr.
  3. Pick one of the eight voices and a speed between 0.5 and 1.5.
  4. Press Generate. Long text is split on sentence boundaries into chunks of at most 400 characters and synthesised one at a time, then joined.
  5. Play it in the inline audio element, or press Download WAV.

What are the eight voices?

Bella, Jasper, Luna, Bruno, Rosie, Hugo, Kiki and Leo — four female and four male style embeddings, loaded from the model's voices.npz pack. Each carries its own natural pacing, so the speed slider is applied on top of a per-voice prior rather than as a raw multiplier. Text is phonemised as US English before inference; other languages will be pronounced as though they were English.

What file do I get?

A standard WAV, saved as tts.wav — 16-bit PCM, mono, 24 kHz, which plays anywhere and imports cleanly into any editor:

tts.wav   RIFF/WAVE, PCM 16-bit, 1 channel, 24000 Hz

Downloading only exists on the KittenTTS side. The browser engine speaks through the OS audio path and never hands JavaScript a buffer, so there is nothing to save from it.

Do I have to download the model every time?

No. The weights and voice pack are stored through the browser's Cache API under a versioned key, so a return visit loads them from disk and the progress bar finishes almost immediately. The cache is skipped in contexts where it is unavailable — some private-browsing modes — in which case you pay the download again.

Does my text get uploaded?

The KittenTTS path never sends your text anywhere: the only network traffic is fetching the model itself, after which generation is pure local CPU work and you can go offline. The browser engine is a thin wrapper over your system's own speech service, and a handful of platform voices are cloud-backed by the OS vendor — if that matters for the text you are reading, use the local model or pick a voice your OS marks as offline.

Related: word count to estimate how long a script will run, and Markdown preview for stripping formatting before you paste.