The local AI toolkit

Local AI software

Find the right app, server or engine. Compare what it does, where it runs and what your hardware needs.

Browse software

Search by name or capability, then narrow the list.

How we compare →

Showing 13 of 13 tools

LocalClaw

Featured

Install and manage OpenClaw from a native Mac app. Your chosen provider or local runtime handles model inference.

Type / useComplete stack

OpenClaw setup · Agent management

System / requirementsmacOS

Apple Silicon only · macOS 13+ for the app, macOS 14+ for local models with LM Studio. Paid app license required.

LM Studio

Easy start

Find, download and chat with local models in a visual app, or expose them through a local API.

Type / useDesktop app + server

Visual chat · Documents · Local API

System / requirementsmacOS · Windows · Linux

Mac: Apple Silicon, macOS 14+. Windows and Linux: supported x64 / ARM builds.

Download, run and serve models with a command-line workflow and local API for other apps.

Type / useModel runner + server

CLI chat · Model management · API

System / requirementsmacOS · Windows · Linux

Hardware acceleration depends on the OS and GPU. Choose a local model for on-device inference.

A desktop workspace for local model chat, file attachments and an OpenAI-compatible server. Cloud providers are optional.

Type / useDesktop app + server

Visual chat · Files · Local API

System / requirementsmacOS · Windows · Linux

Desktop installers for all three systems. The selected model and inference backend determine hardware needs.

Run, train and serve models from a visual workspace. Training and inference have different hardware requirements.

Type / useDesktop app + training

Chat · Fine-tuning · Model serving

System / requirementsmacOS · Windows · Linux

Mac download: Apple Silicon. Training support varies by GPU, system and backend.

Organize documents, chat with your files and run agent workflows. Connect local or cloud model providers.

Type / useDesktop / web workspace

Document chat · RAG · Agents

System / requirementsmacOS · Windows · Linux

Desktop for personal use; Docker for a shared browser workspace. Configure local providers to keep inference local.

A browser interface for model chat and document search. Connect it to Ollama or another supported model server.

Type / useSelf-hosted web interface

Browser chat · Documents · RAG

System / requirementsmacOS · Windows · Linux

Install with Docker or Python. Local inference requires a configured local model backend.

Colibri

Advanced

Streams MoE experts between disk, RAM and optional GPU memory. A research-oriented runtime for large local models.

Type / useStreaming inference engine

MoE offloading · CLI · Local API

System / requirementsmacOS · Windows · Linux

Model-specific setup: GLM-5.2 weights alone need about 372 GB of disk. Speed depends heavily on storage and cache.

Run GGUF models with direct control over CPU and GPU inference, or serve them to other apps.

Type / useInference engine + server

GGUF models · CLI · Local API

System / requirementsmacOS · Windows · Linux

CPU, Apple Metal, CUDA, Vulkan and other backends. Choose the build for your hardware.

A framework for building and running machine learning workloads. Use companion tools such as MLX LM for language models.

Type / useMachine learning framework

Inference · Training · Python / C++

System / requirementsmacOS · Linux

Mac: Apple Silicon. Linux: CUDA or CPU packages. Framework and model-tool support can differ.

FreeToken

NVIDIA

An MoE inference and serving engine that shares work across CPU, memory and NVIDIA GPUs. A desktop app is also available.

Type / useMoE engine + desktop app

Expert caching · Chat · Local API

System / requirementsWindows · Linux

Linux CLI: x86_64, NVIDIA and CUDA 13. Windows uses the desktop setup; check its requirements separately.

Serve language models for concurrent requests and developer applications. Installation depends on the compute backend.

Type / useInference engine + server

Batched inference · Model APIs

System / requirementsLinux · macOS · Windows

Linux runtime; Windows via WSL. Apple Silicon uses the community vLLM-Metal plugin, not the standard Linux package.

An inference runtime for serving models and structured generation workloads, with platform-specific backends.

Type / useInference engine + server

Structured generation · Model APIs

System / requirementsLinux · macOS

Linux CPU / GPU setups. Mac requires Apple Silicon and the documented MLX installation path.

Some tools have several roles. System matches include the documented setup paths shown above; they do not guarantee that a model will fit your RAM or GPU.

Sources checked

Where should you start?

For a visual local chat app, start with LM Studio or Jan. For a command-line model runner and API, look at Ollama. For a document workspace, compare AnythingLLM and Open WebUI.

Need more control?

llama.cpp exposes inference controls directly. Colibri explores disk and memory streaming for large MoE models. Unsloth and MLX also support training workflows. Check each project's hardware notes before choosing.

For OpenClaw setup and ongoing management on a Mac, LocalClaw provides the control center around your chosen provider.

How we compare

This is a curated directory, not a speed ranking or an exhaustive list. Each entry identifies its role, documented systems, main uses and setup limits. The Docs link points to an official source.

System filters include native installers, Docker or Python setups, and clearly labeled alternatives such as WSL or the community vLLM-Metal plugin. Support can differ by backend, GPU, model format and release.

To choose a model that fits your machine, use the AI Index and the RAM/GPU guide. This directory does not automatically read your saved machines.