Local ASR model

A.X K2 Raon-Speech 21B-A3B

Catalogue summary: Bilingual English/Korean speech-language model for local STT, TTS, SpeechQA, SpokenQA and turn-based multimodal chat. The MoE checkpoint has 21.2B total parameters with about 3.5B active, includes KRAFTON's AuT speech encoder and Mimi-style codec, and loads through official Transformers remote-code pipeline examples.

Repository editorial metadata; verify comparative claims in the linked upstream material.

GPU recommendedspeech-to-text transcription2 languagesCC-BY-NC-4.0
Choose an app

Start with Recommended. No terminal commands are shown.

Compare speech models

Desktop app links require the app to be installed. If nothing opens, LocalClaw will show app-download and model-file fallbacks.

Catalogue quality
9.4/10
Catalogue speed
6.4/10
Model size
42.4 GB
Voices
Direct TTS, optional speaker reference conditioning and TTS continuation

Can A.X K2 Raon-Speech 21B-A3B run locally?

A.X K2 Raon-Speech 21B-A3B can run locally for offline speech-to-text. Use the verified setup options on this page; no terminal command is required to choose the right path.

CC-BY-NC-4.0 license. Review upstream restrictions before commercial use.

cloningmultilingualdialoguecontrollable

Audio profile

Cat. quality
9.4
Cat. speed
6.4
Audio
8.4

Best fit

A.X K2 Raon-Speech 21B-A3B is best for offline transcription, speech indexing and local voice pipelines.

Hardware: gpu

Model details

Type
Local ASR model
Family
raon
Latency
medium
Formats
transformerssafetensors
Languages
en, ko
Context
21.2B total / 3.5B active MoE, ~42.4GB BF16 weights; 80GB GPU or suitable multi-GPU setup recommended

Install locally

01
Check runtimeConfirm the backend supports transformers, safetensors on your machine.
02
Open recommended setupUse the app and model links above. LocalClaw does not expose a terminal command.
03
Test locallyRun a short private audio prompt before moving into production workflows.

Good for

  • speech-to-text transcription
  • GPU recommended local workflows
  • cloning, multilingual, dialogue

Watch before shipping

  • Validate pronunciation, latency and artifacts with your own voice samples.
  • Review the upstream license and acceptable-use notes.
  • Benchmark on your target CPU, Apple Silicon or GPU setup.

Related TTS and speech models

CompareBrowse all TTS models Local AIBrowse LLM models macOS appGet LocalClaw