LocalClaw
FeaturedInstall and manage OpenClaw from a native Mac app. Your chosen provider or local runtime handles model inference.
Apple Silicon only · macOS 13+ for the app, macOS 14+ for local models with LM Studio. Paid app license required.
Find the right app, server or engine. Compare what it does, where it runs and what your hardware needs.
Search by name or capability, then narrow the list.
Install and manage OpenClaw from a native Mac app. Your chosen provider or local runtime handles model inference.
Apple Silicon only · macOS 13+ for the app, macOS 14+ for local models with LM Studio. Paid app license required.
Find, download and chat with local models in a visual app, or expose them through a local API.
Mac: Apple Silicon, macOS 14+. Windows and Linux: supported x64 / ARM builds.
Download, run and serve models with a command-line workflow and local API for other apps.
Hardware acceleration depends on the OS and GPU. Choose a local model for on-device inference.
A desktop workspace for local model chat, file attachments and an OpenAI-compatible server. Cloud providers are optional.
Desktop installers for all three systems. The selected model and inference backend determine hardware needs.
Run, train and serve models from a visual workspace. Training and inference have different hardware requirements.
Mac download: Apple Silicon. Training support varies by GPU, system and backend.
Organize documents, chat with your files and run agent workflows. Connect local or cloud model providers.
Desktop for personal use; Docker for a shared browser workspace. Configure local providers to keep inference local.
A browser interface for model chat and document search. Connect it to Ollama or another supported model server.
Install with Docker or Python. Local inference requires a configured local model backend.
Streams MoE experts between disk, RAM and optional GPU memory. A research-oriented runtime for large local models.
Model-specific setup: GLM-5.2 weights alone need about 372 GB of disk. Speed depends heavily on storage and cache.
Run GGUF models with direct control over CPU and GPU inference, or serve them to other apps.
CPU, Apple Metal, CUDA, Vulkan and other backends. Choose the build for your hardware.
A framework for building and running machine learning workloads. Use companion tools such as MLX LM for language models.
Mac: Apple Silicon. Linux: CUDA or CPU packages. Framework and model-tool support can differ.
An MoE inference and serving engine that shares work across CPU, memory and NVIDIA GPUs. A desktop app is also available.
Linux CLI: x86_64, NVIDIA and CUDA 13. Windows uses the desktop setup; check its requirements separately.
Serve language models for concurrent requests and developer applications. Installation depends on the compute backend.
Linux runtime; Windows via WSL. Apple Silicon uses the community vLLM-Metal plugin, not the standard Linux package.
An inference runtime for serving models and structured generation workloads, with platform-specific backends.
Linux CPU / GPU setups. Mac requires Apple Silicon and the documented MLX installation path.
Try a different name, system or use case.
Some tools have several roles. System matches include the documented setup paths shown above; they do not guarantee that a model will fit your RAM or GPU.
Sources checkedFor a visual local chat app, start with LM Studio or Jan. For a command-line model runner and API, look at Ollama. For a document workspace, compare AnythingLLM and Open WebUI.
llama.cpp exposes inference controls directly. Colibri explores disk and memory streaming for large MoE models. Unsloth and MLX also support training workflows. Check each project's hardware notes before choosing.
For OpenClaw setup and ongoing management on a Mac, LocalClaw provides the control center around your chosen provider.
This is a curated directory, not a speed ranking or an exhaustive list. Each entry identifies its role, documented systems, main uses and setup limits. The Docs link points to an official source.
System filters include native installers, Docker or Python setups, and clearly labeled alternatives such as WSL or the community vLLM-Metal plugin. Support can differ by backend, GPU, model format and release.
To choose a model that fits your machine, use the AI Index and the RAM/GPU guide. This directory does not automatically read your saved machines.