Text to Speech MagpieTTS Multilingual Demo

MagpieTTS is NVIDIA's state-of-the-art multilingual text-to-speech system built on the NeMo Speech framework. It employs a two-stage pipeline architecture: a language model generates discrete acoustic tokens from text, which are then decoded into high-fidelity audio using a neural audio codec: NanoCodec. The model synthesizes speech in 5 different English speakers - Sofia, Aria, Jason, Leo, John Van Stan across 12 languages: Arabic, Chinese, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Spanish, and Vietnamese. The model predicts discrete audio codec tokens autoregressively using a transformer encoder-decoder architecture. It employs multi-codebook prediction (typically 8 codebooks) with frame stacking (factor = 2) and a local transformer for fine-grained refinement of high-fidelity audio generation, and leverages techniques like attention priors, classifier-free guidance (CFG), and Group Relative Policy Optimization (GRPO) for improved alignment.

Key Features of the model

  • Multilingual Support — Synthesizes natural speech in 12 languages: Arabic, Chinese, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Spanish, and Vietnamese.
  • Expressive Voices — Multiple voice options with emotional tones and gender variations including 4 proprietary voices and 1 public voice.
  • Text Normalization — Built-in text normalization for handling numbers, abbreviations, and special characters for all languages.

Resources

Note about the demo:

  • Text normalization takes time to load, so if you want faster generation select "Do not apply TN."
  • Text normalization works for all supported languages.
  • Text normalization is required for numbers to be processed and spoken.
  • When the model runs on ZeroGPU Hardware, expect slower generation because the model checkpoint is loaded everytime.
  • As the model's speakers are all native English speakers, expect accented speech in the other languages.
  • The current model can generate long-form speech (more than 20 seconds) in English. However, longer generations can cause timeout due to Huggingface timeout limit.
  • Loan words are not supported at this time. For example, English characters in Mandarin will lead to unexpected results.
  • To add custom phone pronunciations for supported languages, replace the word with it's IPA characters surrounded by '|' characters and add a space inbetween each IPA phone. For example, "Hello world from NeMo Text to Speech." could be written as "Hello world from | ˈ n ɛ m o ʊ | Text to Speech." replacing "NeMo" with "| ˈ n ɛ m o ʊ |".
    • Text normalization does not work with custom pronunciations. To mix the two, normalize the input first and then replace the word with the custom pronunciation.
  • While the checkpoint was trained with 4 Arabic dialects, only ar-MSA is recommended for use.
  • For the enterprise offering, see the MagpieTTS NIM which includes additional native voices in the supported languages, emotional speech capabilities, and optimized batch and latency inference pipeline.
Target Language
Select the target language for the speech to be synthesized in
Target Speaker
Select the target speaker whose voice you would like the the speech to be synthesized in
Apply Text Normalization
Select if you want to apply text normalization to the input text