Local ASR model
VibeVoice ASR
Catalogue summary: Open-source multilingual ASR model (speech-to-text), supporting long-form transcription, timestamps, diarization and hotwords.
Repository editorial metadata; verify comparative claims in the linked upstream material.
GPU recommendedspeech-to-text transcription50 languagesMIT
Choose an app
Compare speech modelsStart with Recommended. No terminal commands are shown.
Open model filesRecommended verified starting pointOpen in UnslothLaunches Unsloth Desktop on this modelOpen Python packageVerified package page without exposing a commandOpen on Hugging FaceFiles, licence and model cardCheck my machineOptional: verify the fit with your saved hardware
Desktop app links require the app to be installed. If nothing opens, LocalClaw will show app-download and model-file fallbacks.
Catalogue quality
9.3/10
Catalogue speed
7.5/10
Model size
14 GB
Voices
N/A (ASR: outputs text)
Can VibeVoice ASR run locally?
VibeVoice ASR can run locally for offline speech-to-text. Use the verified setup options on this page; no terminal command is required to choose the right path.
MIT license. Still verify upstream usage notes before shipping.
streamingrealtimemultilingualdialogue
Audio profile
Best fit
VibeVoice ASR is best for offline transcription, speech indexing and local voice pipelines.
Hardware: gpuapple
Model details
Type
Local ASR model
Family
vibevoice
Latency
medium
Formats
pytorchsafetensors
Languages
en, zh, es, pt, de, ja, ko, fr, ru, id, sv, it, he, nl, pl, tr, th, ar, hi, fi, el, ro, vi, uk
Context
60-minute single-pass ASR with diarization + timestamps
Install locally
01
Check runtimeConfirm the backend supports pytorch, safetensors on your machine.02
Open recommended setupUse the app and model links above. LocalClaw does not expose a terminal command.03
Test locallyRun a short private audio prompt before moving into production workflows.Good for
- speech-to-text transcription
- GPU recommended local workflows
- streaming, realtime, multilingual
Watch before shipping
- Validate pronunciation, latency and artifacts with your own voice samples.
- Review the upstream license and acceptable-use notes.
- Benchmark on your target CPU, Apple Silicon or GPU setup.
Related TTS and speech models
Microsoft Research
VibeVoice 1.5B
Local TTS model · Q 9.4 · Speed 6.5
Microsoft Research
VibeVoice Realtime 0.5B
Local TTS model · Q 9.1 · Speed 9.2
NVIDIA
Parakeet TDT 0.6B v3
Local ASR model · Q 9.5 · Speed 9.7
NVIDIA
Nemotron 3.5 ASR Streaming 0.6B
Local ASR model · Q 9.4 · Speed 9.3
Alibaba Cloud (Qwen Team)
Qwen3-ASR
Local ASR model · Q 9.5 · Speed 9
Mistral AI
Voxtral Mini 4B Realtime 2602
Local ASR model · Q 9.2 · Speed 9.4
OpenAI
Whisper v3 Turbo
Local ASR model · Q 9.1 · Speed 9.5
Audar AI Labs
Audar-ASR-V1-Flash
Local ASR model · Q 9.1 · Speed 9.4