Video catalogue · Verified local path

FastH3 Preview v1 local guide

Open-weight four-forward MiniMax H3 distillation for synchronized text-to-video-and-audio generation.

Choose an app

Start with Recommended. No terminal commands are shown.

Compare video models

Desktop app links require the app to be installed. If nothing opens, LocalClaw will show app-download and model-file fallbacks.

What it does

Text To Audio VideoText To VideoAudio Video GenerationAnimation

The official FastVideo release publishes full safetensors weights, a matching LoRA and a basic_fasth3.py runner. The card defaults to four B200 GPUs for the VSA-H3 CUDA path, while the project documents newer Apple Silicon MLX and DGX Spark recipes; LocalClaw records 128 GB RAM and a 48 GB NVIDIA VRAM workstation floor until smaller RTX/NVFP4 recipes mature.

Hardware figures are practical entry floors, not performance guarantees. Resolution, duration, precision, offloading and runtime versions can materially change memory use.

Strengths

  • Four transformer forwards for text-to-audio-video generation
  • Official FastVideo weights and inference contract
  • Apple Silicon MLX and NVIDIA CUDA recipe coverage

Limits to know

  • Preview checkpoint only supports text-conditioned audio-video, not FL2VA or Ref2VA
  • MiniMax community license carries territory and acceptable-use restrictions
  • Documented high-performance defaults use multi-GPU Blackwell systems