Capability matrix

Match language and voice requirements to an available processing path.

The catalogue below describes declared support, typical hardware paths, and model licence constraints. It is a selection aid, not a benchmark.

Find a compatible local starting point

This filters declared language and hardware paths. It is not a performance benchmark or a guarantee that every voice has equal quality.

Compatible catalogue entries

Kokoro 82M

A lightweight starting point with built-in voices.

CPU and CUDA are established paths; ROCm and Apple Silicon support are experimental upstream paths.

Apache-2.0

Qwen3 TTS

Built-in speakers with CustomVoice, or reference cloning with Base models.

CPU, CUDA, Vulkan, and Metal variants are available; model size and quantization remain selectable.

Apache-2.0

Silero

Efficient built-in voices, including East European and regional packs.

Runs on CPU on Windows and Linux.

MIT or CC BY-NC-SA 4.0, depending on the voice pack

Magpie 357M

Five built-in speakers shared across nine supported languages.

CPU and NVIDIA CUDA variants are available; the first NeMo installation is comparatively large.

NVIDIA Open Model License

The installer remains authoritative for downloadable variants, model size, host compatibility, and the licence shown before installation.

Local speech services

Local speech services

Model size, quantization, input length, and concurrent services all affect practical requirements.

ServiceVoice typeTypical hardware pathDeclared languagesModel licence
Kokoro 82M A lightweight starting point with built-in voices.Ready-madeCPU only · NVIDIA GPU · AMD / Intel / Vulkan-capable GPU · Apple SiliconChinese, English, French, Hindi, Italian, Japanese, Portuguese, Spanish Apache-2.0 Commercial use is permitted under the model terms.
Qwen3 TTS Built-in speakers with CustomVoice, or reference cloning with Base models.Ready-made · CloningCPU only · NVIDIA GPU · AMD / Intel / Vulkan-capable GPU · Apple SiliconChinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian
+1 languages Spanish
Apache-2.0 Commercial use is permitted under the model terms.
XTTS v2 A mature multilingual cloning model using a short reference recording.CloningCPU only · NVIDIA GPUArabic, Chinese, Czech, Dutch, English, French, German, Hungarian, Italian
+7 languages Japanese, Korean, Polish, Portuguese, Russian, Spanish, Turkish
Coqui Public Model License 1.0.0 The model and its outputs are restricted to non-commercial use.
VoxCPM2 A comparatively large multilingual model conditioned by a reference voice.CloningNVIDIA GPUArabic, Burmese, Chinese, Danish, Dutch, English, Finnish, French, German
+21 languages Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Khmer, Korean, Lao, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Tagalog, Thai, Turkish, Vietnamese
Apache-2.0 Commercial use is permitted under the model terms.
Fish S2 Pro Broad declared language coverage with selectable native backends and quantization.CloningCPU only · NVIDIA GPU · AMD / Intel / Vulkan-capable GPU · Apple SiliconAfrikaans, Albanian, Amharic, Arabic, Armenian, Assamese, Azerbaijani, Basque, Belarusian
+74 languages Bengali, Bosnian, Breton, Bulgarian, Burmese, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Faroese, Finnish, French, Galician, Georgian, German, Greek, Gujarati, Haitian Creole, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Javanese, Kannada, Kazakh, Khmer, Korean, Latin, Latvian, Lithuanian, Malay, Malayalam, Maori, Marathi, Mongolian, Nepali, Norwegian, Norwegian Nynorsk, Pashto, Persian, Polish, Portuguese, Punjabi, Romanian, Russian, Sanskrit, Serbian, Shona, Sindhi, Sinhala, Slovak, South Slavey, Spanish, Swahili, Swedish, Tagalog, Tamil, Telugu, Thai, Tibetan, Turkish, Ukrainian, Urdu, Vietnamese, Welsh, Yiddish, Yoruba
Fish Audio Research License Research and non-commercial use; commercial use requires a separate licence.
Voxtral 4B Preset voices through a WGPU-compatible accelerator; no packaged CPU path.Ready-madeNVIDIA GPU · AMD / Intel / Vulkan-capable GPUArabic, Dutch, English, French, German, Hindi, Italian, Portuguese, Spanish CC BY-NC 4.0 Non-commercial use only under the model terms.
Silero Efficient built-in voices, including East European and regional packs.Ready-madeCPU onlyArmenian, Azerbaijani, Bashkir, Belarusian, Bengali, Chuvash, English, Erzya, French
+24 languages Georgian, German, Gujarati, Hindi, Kannada, Kabardian-Cherkess, Kalmyk, Kazakh, Khakas, Kyrgyz, Malayalam, Manipuri, Moksha, Rajasthani, Russian, Spanish, Tajik, Tamil, Tatar, Telugu, Udmurt, Ukrainian, Uzbek, Yakut
MIT or CC BY-NC-SA 4.0, depending on the voice pack The installer shows the applicable pack licence before installation.
Chatterbox Expressive cloning with English and multilingual model choices.CloningCPU only · NVIDIA GPUArabic, Chinese, Danish, Dutch, English, Finnish, French, German, Greek
+14 languages Hebrew, Hindi, Italian, Japanese, Korean, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Turkish
MIT Commercial use is permitted under the model terms.
Magpie 357M Five built-in speakers shared across nine supported languages.Ready-madeCPU only · NVIDIA GPUChinese, English, French, German, Hindi, Italian, Japanese, Spanish, Vietnamese NVIDIA Open Model License The model card marks the checkpoint ready for commercial use under NVIDIA terms.

Speech language overview

Speech language overview

You need only one compatible service. Custom and commercial endpoints may add languages not listed here.

Declared languagesReady-madeCloning
English, French, SpanishKokoro, Qwen3 CustomVoice, Voxtral, Silero, MagpieQwen3 Base, XTTS, VoxCPM2, Fish S2 Pro, Chatterbox
GermanQwen3 CustomVoice, Voxtral, Silero, MagpieQwen3 Base, XTTS, VoxCPM2, Fish S2 Pro, Chatterbox
ItalianKokoro, Qwen3 CustomVoice, Voxtral, MagpieQwen3 Base, XTTS, VoxCPM2, Fish S2 Pro, Chatterbox
PortugueseKokoro, Qwen3 CustomVoice, VoxtralQwen3 Base, XTTS, VoxCPM2, Fish S2 Pro, Chatterbox
RussianQwen3 CustomVoice, SileroQwen3 Base, XTTS, VoxCPM2, Fish S2 Pro, Chatterbox
Chinese, JapaneseKokoro, Qwen3 CustomVoice, MagpieQwen3 Base, XTTS, VoxCPM2, Fish S2 Pro, Chatterbox
KoreanQwen3 CustomVoiceQwen3 Base, XTTS, VoxCPM2, Fish S2 Pro, Chatterbox
HindiKokoro, Voxtral, Silero, MagpieVoxCPM2, Fish S2 Pro, Chatterbox
Arabic, DutchVoxtralXTTS, VoxCPM2, Fish S2 Pro, Chatterbox
Polish, TurkishXTTS, VoxCPM2, Fish S2 Pro, Chatterbox
Czech, HungarianXTTS, Fish S2 Pro
VietnameseMagpieVoxCPM2, Fish S2 Pro
Danish, Finnish, Greek, Hebrew, Malay, Norwegian, Swahili, SwedishVoxCPM2, Fish S2 Pro, Chatterbox
Burmese, Indonesian, Khmer, Tagalog, ThaiVoxCPM2, Fish S2 Pro
LaoVoxCPM2
Additional regional and low-resource languagesSelected Silero packsFish S2 Pro; consult the complete catalogue

Transcription

Transcription

CrispASR exposes both packaged engines with word timestamps, configurable VAD, quantization, and compute backend.

Whisper large-v3

FP16 or Q5_0

100 languages; the broadest packaged option.

Parakeet TDT 0.6B v3

FP16, Q8_0, Q5_0, or Q4_K

25 primarily European languages.

Input and output formats

Input and output formats

Documents
TXT, PDF, EPUB, DOCX, MOBI, pasted text
Subtitles
SRT
Audio sources
AAC, AIFF, FLAC, M4A/MKA, MP3, OGG, Opus, WAV, WMA
Video sources
MP4, MKV, WebM, AVI, MOV
Audio output
M4B, MP3, Opus, FLAC, WAV
Video output
MP4 with selectable audio and subtitle tracks

Catalogue snapshot reviewed 25 July 2026. Update this site-owned snapshot when the Pandrator installer catalogue changes. 2026-07-25