MOC: Which Open-Source TTS (Voice) Model to Use

Map of Content for picking a local, open-weights text-to-speech model as of October 2026. The short answer is a table, sorted by model size. For the long list, see the appendix.

Quick picks

★ marks the pick for each size. Each name links to the model's official page, and the Weights column links to where you download it.

SizeModelWeightsNotes
≤ ~120M (CPU / edge)★ Chatterbox-Nano (Resemble AI)HF110M, MIT. ~3× real-time on 8 CPU cores, voice cloning, tags such as [laugh].
KokoroHF82M, Apache 2.0. Very fast and very popular.
Supertonic 3 (Supertone)HF99M. The repo was archived in July 2026 and gets no further updates.
~120M–800M★ Qwen3-TTS (Alibaba Qwen)HF0.6B (1.7B also available), Apache 2.0. Preset voices, voice design and cloning.
Chatterbox (Resemble AI)GitHubMIT. Emotion control and cloning.
Gepard (nineninesix.ai)HF~0.5B, Apache 2.0. Use it if you need streaming: it's built for real-time voice agents, with first audio in ~50 ms.
OmniVoice (k2-fsa / Xiaomi)HFUse it if you need more languages (600+). It tops the tts-bench cloning votes but can drop words.
~2B★ VoxCPM2 (OpenBMB)HF2B, Apache 2.0, 30 languages, 48 kHz output. Can design a voice from a text description.
Largest / best quality★ Fish Audio S2 ProHFThe Fish Audio Research License only allows research and non-commercial use. Commercial use needs a separate licence.
Not recommendedHiggs TTS 3 (Boson AI)HF4B, research/non-commercial licence.

Older models such as Bark and XTTS are no longer the standard.

The recommendation behind the table

The table is built on a comment from a self-described TTS specialist in r/LocalLLM, What is the best open-source TTS model right now? (posted around July 2026). The comment is tidied here. Links and licence notes were added from the model cards.

I've trained and fine-tuned TTS models for a while and tested every major one.

  • ~120M and below: Chatterbox-Nano, Kokoro, Supertonic. Chatterbox-Nano is the one to go for.
  • ~120M to 800M: Qwen3-TTS-0.6B, Chatterbox, Gepard (if you need streaming). If I had to pick one, Qwen3-TTS, or OmniVoice for more languages.
  • Above that: VoxCPM2 (~2B) if you want something smaller. If larger, Fish Audio S2 Pro. (I wouldn't recommend Higgs Audio TTS 3.)

The gold standard has moved far beyond older models like Bark and XTTS. Lots of people recommend Kokoro, and it's an amazing model, but Chatterbox-Nano outclasses it on basically everything.

This is one practitioner's view, not a benchmark. To hear the models yourself, use the tts-bench listening pages and blind A/B arena below.

Appendix: tts-bench

tts-bench

tts-bench (MIT) benchmarks 75 local TTS models on Windows, Linux and macOS. It measures speed (time to first audio, real-time factor and memory, on CPU, CUDA and Apple Silicon). It also scores each model on naturalness (UTMOS), intelligibility (WER) and speaker similarity (SIM), and you can listen to every model on every prompt. A blind A/B arena ranks models by human votes. All models get the same plain prompts, so expressive controls such as emotion tags are not tested.

It is very thorough and also overwhelming, so use it to check a shortlist rather than to choose from scratch. Snapshot as of 2026-10-08:

Fastest (June 2026 results)

CategoryRigModelWarm TTFARTFx
Fastest CPURyzen 9 9950X3DPiper107 ms59×
Fastest CUDARTX 5090Kokoro67 ms104×
Fastest Apple SiliconM4, 16 GBPiper208 ms32×

Best voice cloning (blind A/B preference)

RankModelW-L-TNote
1OmniVoice24-1-3Best voice match, but can garble or drop words
2Echo-TTS21-1-6Clean 44.1 kHz output
3IndexTTS-216-2-5Keeps accents well

The full table lists 26 models with built-in voices and 37 zero-shot cloning models, each with size, release date and licence.