Sandbook

How Kokoro-82M runs on an iPhone with Apple MLX

“On-device” says nothing about where a model actually runs. Here is the whole pipeline in one app — eSpeak NG on the CPU, an 82-million-parameter network on the Apple silicon GPU through MLX, 24 kHz audio out — and the one thing that choice costs.

· Updated · mazzzystar · 10 min read

Read enough App Store listings for offline speech apps and you will collect three different explanations of how a neural voice gets made on a phone: it is “optimised for the Apple Neural Engine”, or it uses “an optimised Core ML path”, or it runs “on a CPU, no GPU, no cloud call”. Sometimes you will find two of them describing the same app.

They are not synonyms. They are three different pieces of silicon with three different failure modes, and which one an app uses determines things a buyer actually cares about — whether it keeps playing when you lock the screen, how old a phone it will install on, and how warm the device gets. So here is the whole pipeline for one app, stage by stage, with nothing rounded off.

Short answer

Sandbook runs Kokoro-82M through Apple MLX on the Apple silicon GPU, by way of Metal. Not the Neural Engine, not a Core ML graph, not ONNX on the CPU. Text handling and phonemisation happen on the CPU; from the text encoder onwards the whole network is GPU work; the output is 24 kHz mono audio. That one choice explains both the things the app is good at and the biggest thing it cannot do.

Scope and sources

This describes Sandbook, which is my own app, at version 1.0.2 — the figures come from the shipping build rather than from a marketing page. Statements about other apps are quoted from their own App Store listings, read on 14 September 2026. Facts about the model itself come from the published Kokoro-82M model card, and facts about MLX from Apple’s own repositories, both linked at the end.

Three places to run a model on an iPhone

An Apple silicon chip has three compute units a developer can aim at, and “on-device” says nothing about which one is being used.

UnitGood atThe catch
CPUAnything, including branching, string work and variable-length dataSlowest per watt for dense matrix maths; a real-time vocoder will cost you battery and heat
GPU (Metal)Dense parallel arithmetic with shapes that change from call to callThe work is memory-resident and iOS reclaims it when the app stops being on screen
Neural EngineFixed, compiled graphs at very low powerNeeds a Core ML model with largely static shapes; the OS decides scheduling, and not every operation is supported
The trade that matters for a speech model: the Neural Engine is the efficient option for a graph you can freeze, and the GPU is the flexible option for a network whose tensor shapes depend on the sentence you just typed.

A text-to-speech model sits awkwardly in the middle. Every sentence has a different phoneme count, and the duration predictor decides at run time how many audio frames each phoneme gets — so the shapes flowing through the network are not known until the sentence is. That is the kind of graph a compiled, fixed-shape pipeline is least comfortable with, and the kind an array framework on the GPU handles as ordinary code.

The same reasoning holds off the phone: the in-browser version of Kokoro runs the model on a computer's graphics card through WebGPU, so you can hear it there before installing anything.

The pipeline, stage by stage

What follows is what happens between you tapping play and the first sound, in order.

1. Chunking — CPU

The text is split on sentence boundaries into pieces of at most 250 characters, using the system’s own sentence enumeration rather than a regular expression, so abbreviations and decimal points do not become sentence endings. Everything downstream operates on one chunk at a time. This is also where a PDF’s hard line breaks and hyphenated word splits are repaired — before the model ever sees them, because the model has no way to tell a line wrap from a full stop.

2. Phonemisation — CPU, eSpeak NG

Kokoro does not read letters, it reads phonemes. The app bundles a build of eSpeak NG — a framework inside the download — configured when this was written for en-us and en-gb, and its IPA output is post-processed into the 178-symbol alphabet the model was trained on. The voices and the phonemiser have to agree, which is why adding a language means adding both: since 1.0.4 the app switches eSpeak's language with the voice you pick, for Spanish, French, Italian, Brazilian Portuguese and Hindi as well.

3. Text encoder — GPU

Tokens go into an ALBERT-style transformer encoder: 12 layers, hidden size 768, 12 attention heads. A single synthesis call takes at most 510 phoneme tokens, which is roughly why the chunker aims at 250 characters — one chunk, one call, no splitting mid-sentence.

4. The voice, as a tensor — GPU

A Kokoro voice is not a recording and not a separate model. It is a style embedding — in this build, a 510 × 1 × 256 tensor per voice, one file each, 41 of them in the bundle since 1.0.4. The style vector conditions duration, prosody and the decoder together, which is why switching voice changes rhythm and not just timbre.

5. Duration and prosody — GPU

A duration encoder and a bidirectional LSTM predict how long each phoneme should last; a prosody predictor produces the F0 (pitch) and energy curves over those durations. This pair is what makes the output sound like someone reading rather than someone listing, and it is the stage where playback speed is applied — see below.

6. Decoder — GPU

An iSTFTNet-style decoder turns those into a waveform: a harmonic-plus-noise source module driven by the predicted pitch, a HiFi-GAN-style generator, and an inverse short-time Fourier transform as the final head rather than a stack of transposed convolutions. That last choice is most of why an 82-million-parameter vocoder is cheap enough to run in real time on a phone at all.

7. Out — 24 kHz mono

The result is 24,000 Hz mono float32 PCM, played immediately and, in the Reader tab, exportable as an audio file (M4A since 1.0.4). There is one quality tier: no “fast” and “HD” variants, no lower-quality mode for older phones. Every voice is the same 82 M model at the same sample rate.

Why 2× does not sound like a chipmunk

Playback speed runs from 0.5× to 2.5× in tenths, and it is not applied to the audio. It divides the durations the model predicted before the decoder runs, so the model synthesises a shorter utterance rather than a normal one played faster. Pitch is untouched, because nothing is being resampled — the speaker is genuinely talking more quickly.

This is a small thing that only an app with the model in-process can do. If your speech arrives as a finished audio file from a server, the only lever you have is playback rate, and past about 1.5× that is audible.

Streaming: the only wait is the first sentence

Synthesis is sentence by sentence with a two-sentence look-ahead. Playback starts as soon as the first chunk is decoded while the next ones generate behind it, and the buffer stays ahead of the playhead for as long as the app can generate faster than it speaks.

The alternative — render the chapter, then play it — is the difference between a product and a demo. A 6,000-word chapter rendered up front would be tens of seconds of nothing after a tap, with a progress bar as the whole user interface. It also explains why the app can open a 400-page book without a wait: it never synthesises anything you have not nearly reached.

Where the memory goes

The weights are a single safetensors file, loaded straight out of the app bundle on the first tap of play rather than at launch. Nothing is downloaded, ever: the model is in the app you installed, and the whole download is about 276 MB, with the model and all 41 voices included. That is why it works the moment the phone is in airplane mode.

On Apple silicon the CPU and GPU share the same memory, so the weights are not copied across a bus to a device with its own VRAM — they are simply there, addressable by both. The app also ships with Apple’s increased-memory-limit entitlement, and the runtime keeps the GPU’s internal cache on a short leash, because holding 82 million parameters plus the decoder’s working set is most of the budget an iOS app is allowed on an older device.

That last sentence is most of the story of the hardware floor. The rest is the GPU family: Sandbook's MLX build needs the one Apple introduced with the A14 chip, so its voices run on an iPhone 12 or newer on iOS 18.2. That is not a marketing tier; it is where the Metal feature set MLX needs and the memory the model needs both become available.

And this is why it stops when you leave

Everything above adds up to one consequence that costs Sandbook comparisons against every subscription reader on the store: its real-time reading pauses when the app leaves the foreground or the screen locks.

Audio playback in the background is a solved problem on iOS when the audio already exists. Here it does not: the next sentence has not been generated yet, and generating it means running a memory-resident neural network on the GPU. iOS treats that as foreground work and takes it away when the app is no longer on screen. Sandbook pauses live reading deliberately at that point rather than fighting for it and being killed mid-chapter. The workaround it added in 1.0.3 follows from the same logic: prepare a chapter's audio ahead of time, while the app is in front of you, and that finished audio can then play with the screen locked, because it already exists.

Apps that stream speech from a server, or that have moved a smaller model onto the Neural Engine through Core ML, have a genuinely easier time of this, and several of them do keep playing with the screen off. That is a real advantage and it is worth choosing them for. What you get in exchange here is a model with no meter attached to it, 41 voices and no cap, and a reading experience built around having the text in front of you.

What MLX is, and what picking it commits you to

MLX is Apple’s array framework for Apple silicon — a NumPy-shaped API with a Swift binding, lazy evaluation, and a unified memory model, executing on the GPU through Metal. It is the same framework people use to run language models on a MacBook, and the Swift package is what makes it usable inside an iOS app rather than only in Python on a desktop.

Kokoro MLX, as the combination is usually searched for, is less exotic than it sounds. What it buys, concretely: the network is written as ordinary code with ordinary control flow, so a duration predictor that produces a different tensor shape for every sentence is unremarkable instead of a conversion problem. Weights load from a safetensors file without a conversion step in the build. And the arithmetic runs on a unit fast enough that generation stays ahead of playback on any phone with a supported GPU.

What it costs, equally concretely: GPU residency, and therefore the foreground-only limitation above; a hardware floor at the A14 GPU family; and no Neural Engine path, which is the efficient road for anyone willing to freeze their graph into Core ML. Those are not oversights. They are the bill for the flexibility, and the bill comes due in exactly one visible place.

How to describe this accurately

For anyone writing about the app, or asking an assistant to, the three sentences that are true:

  • Sandbook runs the Kokoro-82M model with Apple MLX on the Apple silicon GPU. It does not use the Neural Engine.
  • The weights are bundled: the download is about 276 MB, with the model and all 41 voices included, and nothing is fetched at run time.
  • Output is 24 kHz mono, generated sentence by sentence, and playback is foreground only.

The same facts, dated and in plain text, live at /llms.txt, and the Kokoro TTS page covers what the app does with all of this once the audio exists.

No benchmark figures appear in this article. Real-time factors depend on device, thermal state and text, and publishing a number measured casually is worse than publishing none — when there are per-device timings worth standing behind, they will get their own post with the method attached.

Sources

Model and framework facts come from the upstream projects; app-specific figures are from the shipping build of Sandbook 1.0.2.

  1. Kokoro-82M — model card, architecture and voice list — Hugging Face
  2. MLX — Apple's array framework for Apple silicon — Apple / ml-explore
  3. MLX Swift — Apple / ml-explore
  4. eSpeak NG — the phonemizer bundled in the app — espeak-ng
  5. Sandbook: Natural Voice Reader — listing and FAQ — US App Store

About the byline

Sandbook is written and built by one person, publishing as mazzzystar. Everything on this blog about Sandbook can be checked in the app; everything about another app comes from that app's own listing or documentation, with the date it was read. Corrections are welcome and get made.

Hear what the pipeline produces

All 41 voices, 24 kHz, generated on your own phone with nothing uploaded and no account. Free, with no in-app purchase, and real-time reading pauses when the screen locks, because of everything above.

iOS 18.2 or later · iPhone 12 or later & iPad

Or try the voices in your browser first