Local TTS model

GPT-SoVITS

Catalogue summary: Zero-shot voice cloning TTS combining GPT and SoVITS. Clone any voice from 5 seconds of audio. Extremely popular in the open-source community with 40K+ GitHub stars.

Repository editorial metadata; verify comparative claims in the linked upstream material.

Apple Silicon readytext-to-speech generation5 languagesMIT
Choose an app

Start with Recommended. No terminal commands are shown.

Compare speech models
Catalogue quality
9.1/10
Catalogue speed
7/10
Model size
2 GB
Voices
Zero-shot cloning from 5s

Can GPT-SoVITS run locally?

GPT-SoVITS can generate speech locally for private voice workflows. Use the verified setup options on this page; no terminal command is required to choose the right path.

MIT license. Still verify upstream usage notes before shipping.

cloningmultilingualemotion

Audio profile

Cat. quality
9.1
Cat. speed
7
Audio
8.4

Best fit

GPT-SoVITS is best for local voice cloning and expressive speech generation.

Hardware: gpuapple

Model details

Type
Local TTS model
Family
gptsovits
Latency
medium
Formats
pytorch
Languages
en, zh, ja, ko, yue
Context
GPT + SoVITS hybrid

Install locally

01
Check runtimeConfirm the backend supports pytorch on your machine.
02
Open recommended setupUse the app and model links above. LocalClaw does not expose a terminal command.
03
Test locallyRun a short private audio prompt before moving into production workflows.

Good for

  • text-to-speech generation
  • Apple Silicon ready local workflows
  • cloning, multilingual, emotion

Watch before shipping

  • Validate pronunciation, latency and artifacts with your own voice samples.
  • Review the upstream license and acceptable-use notes.
  • Benchmark on your target CPU, Apple Silicon or GPU setup.

Related TTS and speech models

CompareBrowse all TTS models Local AIBrowse LLM models macOS appGet LocalClaw