Local TTS model

MagpieTTS Multilingual 357M

Catalogue summary: NVIDIA's 364M parameter multilingual neural TTS model for local speech synthesis across 12 languages. The v2607 checkpoint supports fixed speaker voices, long-form generation, IPA pronunciation control and a practical local path through NeMo Speech or NeMo-Speech.cpp with GGUF.

Repository editorial metadata; verify comparative claims in the linked upstream material.

Apple Silicon readytext-to-speech generation12 languagesNVIDIA Open Model License
Choose an app

Start with Recommended. No terminal commands are shown.

Compare speech models

Desktop app links require the app to be installed. If nothing opens, LocalClaw will show app-download and model-file fallbacks.

Catalogue quality
9/10
Catalogue speed
8.6/10
Model size
1.5 GB
Voices
Five fixed speaker voices across 12 languages; zero-shot cloning removed

Can MagpieTTS Multilingual 357M run locally?

MagpieTTS Multilingual 357M can generate speech locally for private voice workflows. Use the verified setup options on this page; no terminal command is required to choose the right path.

NVIDIA Open Model License license. Review upstream restrictions before commercial use.

multilinguallong-formcontrollablelow-latency

Audio profile

Cat. quality
9
Cat. speed
8.6
Audio
8.9

Best fit

MagpieTTS Multilingual 357M is best for fast on-device voice responses and local assistants.

Hardware: gpucpuapple

Model details

Type
Local TTS model
Family
magpie
Latency
low
Formats
nemoggufpytorch
Languages
ar, de, en, es, fr, hi, it, ja, ko, pt, vi, zh
Context
364M params, 1.47GB .nemo and 449MB v2602 F16 GGUF, NeMo Speech plus NeMo-Speech.cpp path

Install locally

01
Check runtimeConfirm the backend supports nemo, gguf, pytorch on your machine.
02
Open recommended setupUse the app and model links above. LocalClaw does not expose a terminal command.
03
Test locallyRun a short private audio prompt before moving into production workflows.

Good for

  • text-to-speech generation
  • Apple Silicon ready local workflows
  • multilingual, long-form, controllable

Watch before shipping

  • Validate pronunciation, latency and artifacts with your own voice samples.
  • Review the upstream license and acceptable-use notes.
  • Benchmark on your target CPU, Apple Silicon or GPU setup.

Related TTS and speech models

CompareBrowse all TTS models Local AIBrowse LLM models macOS appGet LocalClaw