Sandbook

Local TTS · on-device

Local TTS, running on your iPhone

Sandbook is a local TTS app for iPhone and iPad: the Kokoro text-to-speech model runs on the device, so your text is turned into speech without being sent to a server. 41 voices in six languages, airplane mode included, free with no account and no meter.

iOS 18.2 or later · iPhone 12 or later & iPad · No account

In short

What is a good local TTS app for iPhone?

Sandbook runs the Kokoro-82M text-to-speech model on the iPhone's GPU through Apple MLX, so books, PDFs, web pages, camera scans and pasted text are read aloud without any text being sent to a server. It has 41 voices in six languages, highlights each word as it is spoken, and works in airplane mode. It is free, with no in-app purchase, no account and no character limit, and needs iOS 18.2 on an iPhone 12 or later, or an iPad with an A14 or M1 chip or later.

Last verified 29 September 2026

What “local TTS” means

TTS is text to speech: software that turns written words into a spoken voice. Local TTS means the model that does the turning runs on your own device. Cloud TTS means it runs on someone else's computer: your text is sent over the network, a server synthesizes the audio, and the audio is sent back.

Both produce a voice, so from the outside they can look the same. The difference is where the work happens, and three things follow from that. With cloud TTS the text has to travel, a connection is required at the moment of speech, and every sentence costs the provider server time, which is why cloud readers count characters, minutes or hours. With local TTS the text stays where it is, the connection is irrelevant, and there is no running bill to recover.

A few years ago local meant the flat, robotic system voices. That changed when neural speech models became small enough to fit on a phone and fast enough to run there in real time. Local TTS is now a trade of breadth for independence, not of quality for privacy.

Why local TTS matters on a phone

A phone is where reading happens in the least predictable places, which is exactly where a server dependency hurts most.

  • Privacy by architecture. Your books, PDFs and scanned pages are not uploaded, because nothing needs them anywhere else. Everything stays on your device, and there is no account and no tracking. The long version is on the private text to speech page.
  • Airplane mode, tunnels, bad signal. Speech made on the device does not notice the network is gone. A chapter opened on a flight sounds the same as one opened at home. More on that under offline text to speech.
  • No per-character meter. When the voice costs nothing to produce, there is no reason to ration it: no daily quota, no monthly hours, no premium voice tier.
  • No queue. There is no upload, no rendering job and no waiting for a file to come back. Press play and the first sentence starts as soon as it is generated.

How Sandbook does local TTS

The engine is Kokoro-82M, an open-weight neural text-to-speech model with 82 million parameters. The download is about 276 MB, with the model and all 41 voices included, so nothing is fetched on first launch and there is no API key or activation step.

Sandbook runs the model on the device's GPU through Apple MLX, Apple's framework for machine learning on its own chips. Text is split into sentences and each one is synthesized as you listen, at 24 kHz, so playback starts after the first sentence rather than after the whole chapter. Speed runs from 0.5× to 2.5×, and each word is highlighted as it is spoken.

The same Kokoro model runs locally in a desktop browser, too: try Kokoro TTS in a browser tab, where the speech is generated on your computer and nothing is uploaded.

There are 41 voices in six languages: 28 in English (20 American, 8 British), 3 in Spanish, 1 in French, 2 in Italian, 3 in Brazilian Portuguese and 4 in Hindi. They read EPUB, PDF, TXT, RTF, Markdown and HTML files, web pages, camera scans and photos (recognized on the device), and typed or pasted text. Text in the Reader tab can be exported as an M4A file.

What local TTS costs you, honestly

Running the model on the phone is a choice with a bill, and it is paid in the following:

  • A large download. The download is about 276 MB, with the model and all 41 voices included. Cloud apps are small because their voices live elsewhere.
  • A newer device. Sandbook needs iOS 18.2 on an iPhone 12 or later, or an iPad with an A14 or M1 chip or later. If your iPhone is older, see text to speech on older iPhones.
  • Real-time reading pauses when the screen locks. The model uses the GPU, and iOS suspends that work when the app leaves the foreground. To listen with the screen locked, prepare a chapter's audio first, then lock the phone and play it.
  • Fewer languages than cloud services. Six voice languages, not dozens. A server can hold a catalogue no phone has room for.

How to check it is really local

Any app can say it runs on the device. The test takes a minute and does not require trusting anyone:

  1. Open a book, a PDF or some pasted text in the app.
  2. Turn on airplane mode, with Wi-Fi off as well.
  3. Press play and listen to a full chapter, then try a voice you have not used yet.

Speech generated on a server cannot survive that test. Speech generated on the phone does not notice it. In Sandbook the only thing that waits for a connection is adding a new web page, since the page has to be fetched from the site that hosts it. Scanning a printed page with the camera works in airplane mode too, because the text recognition also runs on the device.

Questions about local TTS

Is local TTS as good as cloud TTS?

For reading a book aloud, the gap is much smaller than it used to be. Sandbook runs Kokoro-82M, a neural text-to-speech model, at 24 kHz on the iPhone's GPU, which is a different class of voice from the older robotic system voices that on-device speech once meant. Where cloud TTS still wins is breadth: a server can hold thousands of voices in dozens of languages, and a phone cannot. You can judge the quality yourself, because every one of Sandbook's 41 voices has a sample on the voices page.

Does local TTS work offline?

Yes, that is what makes it local. In Sandbook the model and all 41 voices ship inside the app, so EPUB, PDF, TXT, RTF, Markdown and HTML files, camera scans and photos, and typed or pasted text all read aloud in airplane mode. The one thing that needs a connection is adding a new web page, because the page has to be fetched before it can be read; once saved, it reads offline too.

Which iPhones can run local TTS?

Sandbook needs iOS 18.2 or later on an iPhone 12 or later, or an iPad with an A14 or M1 chip or later. The floor is set by the model, not by a business decision: generating neural speech in real time needs the GPU and memory of those chips. Older iPhones can still use the text to speech built into iOS, which runs on any supported device.

Is Sandbook's local TTS free?

Yes, completely. There is no in-app purchase, no subscription, no ads, no account and no character or minute limit. Because the speech is generated on your own device, reading a chapter costs nothing to anyone, so there is nothing to meter. All 41 voices in six languages are available from first launch.

Local TTS you can test in airplane mode

Install it, turn the radios off, and read a whole chapter. Everything stays on your device. Free, no account, 41 voices made on your iPhone.

iOS 18.2 or later · iPhone 12 or later & iPad

Or try the voices in your browser first