August 2026 correction: this edition uses the same hardware-fit and use-case rules as the LocalClaw finder. It only recommends models with a credible local artifact and a realistic memory tier. Performance still depends on quantization, context length, runtime and hardware.
How this shortlist works
A model qualifies when its base release is identifiable, its license is documented, and users have a practical local path such as official or mature GGUF/MLX artifacts. We do not rank API-only models, unverified community names or giant checkpoints as laptop models.
- 8GB tier: compact Q4 checkpoints with enough headroom for runtime and context.
- 16GB tier: strong 8-14B models, plus specialized compressed releases.
- 32GB tier: practical 27-35B Q4 models and sparse MoE options.
- 64GB+ tier: 70B Q4-class checkpoints and larger contexts.
The practical August 2026 shortlist
8GB · FAST STARTER
Nanbeige4.2 3B
A compact recent model for low-memory systems. Choose it when responsiveness and footprint matter more than maximum reasoning depth.
8GB · AGENTIC
Agents-A1 4B
A small action-oriented model with a practical local size. Useful for tool experiments and lightweight assistants.
16GB · MULTIMODAL
Gemma 4 12B
Google's Apache 2.0 multimodal sweet spot. The official Q4-class footprint is about 6.7GB before context and runtime overhead.
16GB · GENERAL
Ministral 3 14B Instruct
A balanced mid-size option for chat and instruction following when an 8B model is too limited but 27B is too heavy.
16GB · CUSTOM RUNTIME
Bonsai 27B
A ternary 27B release with an unusually small 7.2GB weight footprint. It is compelling, but its custom execution path is less universal than ordinary GGUF.
32GB · DENSE REASONER
Qwen 3.8 27B
A strong dense Qwen option for reasoning, coding and multilingual work on 32GB-class systems.
32GB · SPARSE MOE
Qwen 3.6 35B-A3B
A sparse alternative with only a fraction of its parameters active per token. Full weights still need memory, but inference can be efficient.
64GB+ · WORKSTATION
Llama 3.3 70B
Meta's official Llama 3.3 release is 70B, not 8B. A Q4 build is a high-memory workstation target rather than a normal laptop recommendation.
Best model class by available memory
| Memory | Good target | Examples |
|---|---|---|
| 8GB | 3B-8B Q4, short-to-medium context | Nanbeige4.2 3B, Agents-A1 4B, Gemma4 E4B |
| 16GB | Compact 4B-14B Q4 models | LFM2.5 8B-A1B, Granite 4.1 8B, Qwen 3 8B |
| 32GB | 12B-32B Q4 models | Gemma 4 12B, Qwen 3 32B, Qwen 3.8 27B |
| 64GB+ | 27B-70B Q4 or smaller models with more context | Qwen 3.8 27B, Qwen 3 Coder 30B, Nemotron 3 70B |
Important models that are not consumer recommendations
Kimi K3, GLM-5.2, Hy3 and MiniMax M3 are worth following for research, coding and agentic progress. Their complete weight sets and practical serving requirements place them in server, cluster or extreme workstation territory. Quantization can reduce storage and memory, but it does not turn every frontier MoE into a sensible 16GB laptop install.
LocalClaw covers these releases in model pages and technical articles while keeping them out of the default consumer shortlist.
How to choose without chasing hype
- Start with available memory, not parameter marketing.
- Reserve headroom for context, KV cache, runtime and the operating system.
- Prefer Q4_K_M or the project's recommended quant for an initial fit test.
- Verify the runtime path. A custom ternary engine, MLX build and ordinary GGUF are not interchangeable.
- Test your own workload. Coding, multilingual chat, vision and long context reward different models.
LocalClaw verdict: there is no universal number-one local model. The best model is the strongest verified checkpoint that fits your hardware with enough context headroom.
Compare all 213 local model records