MOC: Which Open-Source TTS (Voice) Model to Use
MOC: Which Open-Source TTS (Voice) Model to Use
Map of Content for picking a local, open-weights text-to-speech model as of October 2026. The short answer is a table, sorted by model size. For the long list, see the appendix.
Quick picks
★ marks the pick for each size. Each name links to the model's official page, and the Weights column links to where you download it.
| Size | Model | Weights | Notes |
|---|---|---|---|
| ≤ ~120M (CPU / edge) | ★ Chatterbox-Nano (Resemble AI) | HF | 110M, MIT. ~3× real-time on 8 CPU cores, voice cloning, tags such as [laugh]. |
| Kokoro | HF | 82M, Apache 2.0. Very fast and very popular. | |
| Supertonic 3 (Supertone) | HF | 99M. The repo was archived in July 2026 and gets no further updates. | |
| ~120M–800M | ★ Qwen3-TTS (Alibaba Qwen) | HF | 0.6B (1.7B also available), Apache 2.0. Preset voices, voice design and cloning. |
| Chatterbox (Resemble AI) | GitHub | MIT. Emotion control and cloning. | |
| Gepard (nineninesix.ai) | HF | ~0.5B, Apache 2.0. Use it if you need streaming: it's built for real-time voice agents, with first audio in ~50 ms. | |
| OmniVoice (k2-fsa / Xiaomi) | HF | Use it if you need more languages (600+). It tops the tts-bench cloning votes but can drop words. | |
| ~2B | ★ VoxCPM2 (OpenBMB) | HF | 2B, Apache 2.0, 30 languages, 48 kHz output. Can design a voice from a text description. |
| Largest / best quality | ★ Fish Audio S2 Pro | HF | The Fish Audio Research License only allows research and non-commercial use. Commercial use needs a separate licence. |
| Not recommended | Higgs TTS 3 (Boson AI) | HF | 4B, research/non-commercial licence. |
Older models such as Bark and XTTS are no longer the standard.
The recommendation behind the table
The table is built on a comment from a self-described TTS specialist in r/LocalLLM, What is the best open-source TTS model right now? (posted around July 2026). The comment is tidied here. Links and licence notes were added from the model cards.
I've trained and fine-tuned TTS models for a while and tested every major one.
- ~120M and below: Chatterbox-Nano, Kokoro, Supertonic. Chatterbox-Nano is the one to go for.
- ~120M to 800M: Qwen3-TTS-0.6B, Chatterbox, Gepard (if you need streaming). If I had to pick one, Qwen3-TTS, or OmniVoice for more languages.
- Above that: VoxCPM2 (~2B) if you want something smaller. If larger, Fish Audio S2 Pro. (I wouldn't recommend Higgs Audio TTS 3.)
The gold standard has moved far beyond older models like Bark and XTTS. Lots of people recommend Kokoro, and it's an amazing model, but Chatterbox-Nano outclasses it on basically everything.
This is one practitioner's view, not a benchmark. To hear the models yourself, use the tts-bench listening pages and blind A/B arena below.
Appendix: tts-bench
tts-bench (MIT) benchmarks 75 local TTS models on Windows, Linux and macOS. It measures speed (time to first audio, real-time factor and memory, on CPU, CUDA and Apple Silicon). It also scores each model on naturalness (UTMOS), intelligibility (WER) and speaker similarity (SIM), and you can listen to every model on every prompt. A blind A/B arena ranks models by human votes. All models get the same plain prompts, so expressive controls such as emotion tags are not tested.
It is very thorough and also overwhelming, so use it to check a shortlist rather than to choose from scratch. Snapshot as of 2026-10-08:
Fastest (June 2026 results)
| Category | Rig | Model | Warm TTFA | RTFx |
|---|---|---|---|---|
| Fastest CPU | Ryzen 9 9950X3D | Piper | 107 ms | 59× |
| Fastest CUDA | RTX 5090 | Kokoro | 67 ms | 104× |
| Fastest Apple Silicon | M4, 16 GB | Piper | 208 ms | 32× |
Best voice cloning (blind A/B preference)
| Rank | Model | W-L-T | Note |
|---|---|---|---|
| 1 | OmniVoice | 24-1-3 | Best voice match, but can garble or drop words |
| 2 | Echo-TTS | 21-1-6 | Clean 44.1 kHz output |
| 3 | IndexTTS-2 | 16-2-5 | Keeps accents well |
The full table lists 26 models with built-in voices and 37 zero-shot cloning models, each with size, release date and licence.