Best local text-to-speech models

Speech synthesis models you can run yourself. Ranking these alongside transcription models, as a single 'speech' list would, compares two unrelated tasks.

From the file· live filter over real dataFrom the file· 23 models

How this is ranked

Ranked by downloads among synthesis models. Streaming latency matters more than throughput here, and nobody has measured it on consumer hardware — including us.

Best local text-to-speech models

#ModelParamsSmallest quantSmallest
1Ace-Step1.5160M0.04 GiB0.04 GiB
2Qwen3-TTS-12Hz-0.6B-Base915M0.50 GiB0.50 GiB
3OmniVoice613M0.56 GiB0.56 GiB
4orpheus-3b-0.1-ft3.8B0.91 GiB0.91 GiB
5VieNeu-TTS-0.3B244M0.19 GiB0.19 GiB
6neutts-air748M0.43 GiB0.43 GiB
7s2-proMoE4.6B2.40 GiB2.40 GiB
8Fun-CosyVoice3-0.5B-25120.34 GiB0.34 GiB
9VoxCPM22.3B1.57 GiB1.57 GiB
10chatterbox264M0.17 GiB0.17 GiB
11Kokoro-82M82M0.13 GiB0.13 GiB
12VieNeu-TTS553M0.39 GiB0.39 GiB
13csm-1b1.6B1.34 GiB1.34 GiB
14VibeVoice-Realtime-0.5B1.0B0.65 GiB0.65 GiB
15tada-3b-ml4.2B5.20 GiB5.20 GiB
16MOSS-TTS-Local-Transformer-v1.54.6B5.56 GiB5.56 GiB
17Qwen3-TTS-12Hz-1.7B-Base1.9B0.94 GiB0.94 GiB
18orpheus-3b-0.1-pretrained3.8B1.40 GiB1.40 GiB
19ZONOS24.58 GiB4.58 GiB
20Qwen3-TTS-12Hz-1.7B-VoiceDesign1.9B1.90 GiB1.90 GiB
21VibeVoice-1.5B2.7B1.76 GiB1.76 GiB
22VoxCPM-0.5B728M2.45 GiB2.45 GiB
23Qwen3-TTS-12Hz-1.7B-CustomVoice1.9B1.90 GiB1.90 GiB
Spec sheetFrom the fileFrom the filewhat these mean