Sandbook

Kokoro TTS online, free and private in your browser

Kokoro TTS online, free: type text, pick one of 41 Kokoro voices in six languages, and generate speech right here in your browser — nothing is uploaded.

In short

Can I use Kokoro TTS online for free?

Yes. This page runs Kokoro-82M, the open-weight text-to-speech model, inside your browser tab, free and with no account. It offers all 41 voices Sandbook ships, in English, Spanish, French, Italian, Brazilian Portuguese and Hindi, up to 1,500 characters at a time, with a WAV download. Your text is never uploaded, because the speech is generated on your own device.

Last verified 2 October 2026

How it works

The model runs on your own computer, plays the first sentence while the rest renders, and gives you a WAV to download. When you press Start, your browser downloads the Kokoro-82M model once and keeps it in its cache. From then on everything happens in this tab. Your text is split into sentences, each sentence is turned into phonemes, and the model turns the phonemes into 24 kHz audio in the voice you picked. The first sentence starts playing as soon as it is ready, while the next ones render behind it.

English text goes through Kokoro's own English phonemizer, which also spells out numbers, prices and times. Spanish, French, Italian, Portuguese and Hindi go through eSpeak NG, the same phonemizer Kokoro uses for those languages, loaded only when you pick one of those voices. The inference runs in a background worker, so the page stays responsive.

What to expect: WebGPU or CPU

The tool picks the engine for you. There is no quality setting to get wrong, because the smaller quantized models produce garbled speech on WebGPU.

  • WebGPU (graphics card): the full-precision model, a one-time 326 MB download. On a recent laptop, a 40-word paragraph takes about 3 seconds and the first sentence plays after about 1 second. Current Chrome and Edge on a computer have WebGPU; so does Safari 26 on a Mac.
  • CPU (WebAssembly): the 8-bit model, a one-time 92 MB download, used when WebGPU is not available. It is slower than real time: expect the first sentence after 10 to 15 seconds and 20 to 40 seconds for the same paragraph, depending on your processor.

Either way, the second visit skips the download: the model loads from the browser cache in a few seconds. Plan on 1 to 1.5 GB of memory in the tab while it runs.

Six languages, 41 voices

The voice decides the language. The tool offers the same 41 voices as the Sandbook app, with the app's names and descriptions:

  • English: 28 voices, 20 American and 8 British, including Heart, Bella, Michael, Emma and George.
  • Spanish: 3 voices (Dora, Alex, Santa).
  • French: 1 voice (Siwis).
  • Italian: 2 voices (Sara, Nicola).
  • Brazilian Portuguese: 3 voices (Dora, Alex, Santa).
  • Hindi: 4 voices (Alpha, Beta, Omega, Psi).

To audition a voice without loading anything, press the play button next to the voice picker, or listen to all the Kokoro voice samples. Speed runs from 0.5× to 2.0×, applied inside the model, so a fast voice still sounds like a person talking quickly.

Download, usage rights and limits

Download. Every generation can be saved as a 24 kHz mono WAV file. MP3 and M4A export are not available yet.

Usage rights. Kokoro-82M is published under the Apache-2.0 license, which allows commercial use. The audio you generate is yours; Sandbook claims nothing over it.

Limits. Up to 1,500 characters per generation, about a minute and a half of speech. The model needs a desktop-class amount of memory, so the tool is meant for computers. Phones may not have the memory for the browser version, and on a phone the page shows voice samples before offering to load anything. Closing the tab right after the first download can occasionally stop the browser from finishing its cache write, in which case the next visit downloads the model again.

Privacy. Everything stays in your browser. Your text and the audio never leave this tab; the only files the tool fetches are the model, the voice you pick and, for non-English voices, the phonemizer. The page counts how often the tool is used, never what you type.

Want this offline on iPhone?

This page is a demo of the model. Sandbook is the app: it runs the same Kokoro-82M model on iPhone and iPad through Apple MLX, fully offline, with all 41 voices and no length limit. It reads EPUB books and PDFs, web pages and scanned pages aloud, highlights each word as it is spoken, and remembers where you stopped. Free, with no account and nothing to buy. It needs an iPhone 12 or later on iOS 18.2.

How the model runs on a phone is explained on the Kokoro TTS app for iPhone page. If you only need a phone to read the screen aloud, iOS already can: text to speech on iPhone with the built-in settings. And for how Sandbook compares with Speechify, NaturalReader and the other readers, see the iPhone reader alternatives.

iPhone and iPad only · Free, no in-app purchase

Kokoro TTS online questions

Is this Kokoro TTS online tool really free?

Yes. There is no account, no sign-up and no count of generations. Each generation takes up to 1,500 characters, and you can run as many as you like. It costs nothing to run because the model runs on your own computer, not on a server.

Can I use the audio I generate commercially?

Kokoro-82M is released under the Apache-2.0 license, which allows commercial use. Sandbook claims no rights over the audio you make here: download the WAV and use it in a video, a course or a product. The voices are synthetic, so no real person's voice is involved.

Why is it slow on my computer?

Without WebGPU the tool runs a 92 MB model on the CPU through WebAssembly, which is slower than real time: roughly 20 to 40 seconds for a 40-word paragraph on a fast laptop, with the first sentence playing after 10 to 15 seconds. A browser with WebGPU, such as current Chrome or Edge on a computer, uses the full model on the graphics card and does the same paragraph in about 3 seconds.

Does it work on iPhone or Android?

Phones may not have the memory for it. The model needs about 1 GB inside one browser tab, and Safari on iPhone closes tabs that use too much. On a phone the page plays voice samples first and offers a beta button to try anyway. On iPhone and iPad, the free Sandbook app runs the same model natively and offline.

Does it work offline?

Partly. After the first load, the model files are cached in your browser, and speech is generated in the tab with no server involved. Opening the page itself still needs a connection. For fully offline Kokoro TTS on iPhone or iPad, use the Sandbook app, which bundles the model.