A couple of weeks ago I released Chickadee, an open-source Chrome extension that reads web pages aloud. It runs entirely on your own computer and it’s completely free. I built it because I wanted a quick and easy way to have the browser read to me that’s free and that I could trust wouldn’t share my data.
The voice model I used for Chickadee is Kokoro-82M, an open text-to-speech model. The first time I heard it, I was almost taken aback at how natural a model under 100 million parameters could sound. Here it is reading a sentence out loud from the Wikipedia page for chickadees:
Mountain chickadees can hide as many as 80,000 individual seeds each, which they retrieve during the winter.
But running Kokoro inside a browser is still very resource intensive. The model download takes up 310 MB, and it needs WebGPU, which many laptops don’t support well. Even on machines that it can run on, it consumes a lot of resources.
Around the same time as I was releasing Chickadee, I was beginning to grow interested in model distillation, a technique that trains a smaller student model to approximate a larger teacher model’s performance. I wanted to see if I could distill Kokoro into a smaller model that I could ship with Chickadee and that could easily run on any device. Another insight I had is that while Kokoro is able to produce audio for many voices, I could limit the smaller model to just one voice to further save on size. In a way, this is an example of domain-specific distillation.
The result is Paradee, an 8M parameter voice model distilled from Kokoro that sounds like Kokoro’s af_heart voice at about a tenth of the size. Here are both of them reading a line from Wikipedia’s page on black-capped chickadees.
Males and females are generally similar, although males have a larger bib.
As you can hear, they’re not identical. If you put them side by side, Kokoro is still a little cleaner. On its own though, Paradee is easy to listen to and very natural sounding for a model its size.
Here is a quick comparison of Kokoro and Paradee:
| Kokoro-82M | Paradee | |
|---|---|---|
| Parameters | 82M | 8M |
| File size | 325 MB | 9 MB |
| Naturalness score (out of 5) | 4.52 | 4.41 |
| Words misheard | 5.7% | 6.0% |
These are measured on 200 held-out sentences. The naturalness score comes from UTMOS, a model trained on human ratings that predicts how natural a clip sounds. “Words misheard” is how many words Whisper, a speech recognition model, gets wrong when it transcribes the audio.
A brief background on how Kokoro works
Kokoro can largely be split into two halves, a text side and an acoustic side (or decoder). I take advantage of this two component architecture when distilling the student model, but more on that later.

Here is an overview of how Kokoro turns text into speech. First, a library called misaki turns the text into phonemes, the individual sounds of speech. “Males” becomes something like “mˈAlz”. Misaki mostly looks words up in a pronunciation dictionary. It isn’t part of Kokoro’s 82M parameters, and Paradee uses it unchanged. From then on, Kokoro only works with phonemes.
The text side (28M parameters) reads the phonemes and works out how to say them. It decides how long to hold each sound, how the pitch should rise and fall, and how loud each sound should be. It also produces a 512-dimensional feature encoding for each phoneme.
The decoder (53M parameters) turns the timing, pitch, loudness and feature encodings into audio, in two stages. First, each phoneme’s features are stretched to its predicted length, and the decoder blocks combine them with the pitch, loudness and voice into one representation per frame. Then the waveform generator upsamples those frames into 24,000 samples of audio per second. To make voiced sounds easier, Kokoro also precomputes a helper tone from the predicted pitch, using a fixed formula rather than anything learned. The waveform generator uses this tone as an extra input alongside the frames. The waveform generator is a convolutional network that gradually upsamples the frames and predicts a spectrogram. For every frequency at every moment, the spectrogram holds two values: the magnitude is how loud that frequency is, and the phase is where its wave is in its cycle. A fixed inverse Fourier transform then turns that spectrogram into the sound wave.
Here’s Kokoro going through each step for the chickadee sentence. Scroll to follow it through.








