#1 · Best match
35B (3B active, MoE) · 48GB min · Q4_K_M · 21.72GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official MIT 35B MoE reasoning model from Ornith AI with about 3B active parameters, 262K context, strong agentic-coding positioning and official Q4_K_M GGUF plus MLX/Ollama/llama.cpp local paths for larger workstations.
chatcodereasoningagenticlong-context
#2 · Best match
29.8B multimodal · 24GB min · K-Quant 17GB Q4_K_M · 17GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Meta Superintelligence Lab local agent model with text+image input, 131K context, Apache 2.0 weights and official GGUF/ExecuTorch artifacts. The K-Quant 17GB build targets 24GB machines; 32GB is safer for vision and long-context sessions.
chatcodereasoningagenticvision
#3 · Best match
27B · 32GB min · Q4_K_M · 16.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official Qwen dense 27B vision-language release with Apache 2.0 weights, 262K native context, thinking controls and strong agentic coding benchmarks. Practical local path through Unsloth and LM Studio-compatible GGUF artifacts.
chatcodereasoningvisionagentic
#4 · Best match
29.3B · 32GB min · Q4_K_M · 18GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.2 30B brings the permissive Apache 2.0 Granite stack to workstation-class local reasoning, RAG, coding and tool-use workflows with GGUF and MLX community artifacts.
chatcodereasoningtool-callingpower
#5 · Best match
8.8B · 8GB min · Q4_K_M · 5.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.2 8B instruct model with Apache 2.0 weights, 128K context, thinking-mode chat template, tool calling and practical GGUF plus MLX paths for everyday local machines.
chatcodereasoningtool-callingstandard
#6 · Best match
35B (3B active, MoE) · 32GB min · Q4_K_M · 19GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Qwen Team open-weight MoE for agentic coding and multimodal work. 35B total / 3B active, 262K native context, Apache 2.0, and strong GGUF availability through Unsloth and LM Studio-compatible artifacts.
chatcodereasoningvisionagentic
#7 · Best match
32B · 32GB min · Q4_K_M · 20GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Qwen 3 dense 32B open-weight model with hybrid thinking and non-thinking modes, strong reasoning and coding support, and a practical Q4_K_M GGUF path for 32GB-class local machines.
chatcodereasoningpowerquality
#8 · Best match
9B · 16GB min · Q4_K_M · 5.63GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official MIT reasoning model from Ornith AI with a 262K native context window, tool-calling focus, and official Q4_K_M GGUF, MLX and Ollama/llama.cpp paths for local coding-agent experiments on 16GB+ machines.
chatcodereasoningagenticlong-context
#9 · Best match
30B (3B active, MoE) · 48GB min · Q4_0 · 18.9GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: NVIDIA OpenMDW-1.1 hybrid Mamba/MoE/attention model for local agentic inference. The official GGUF path from ggml-org includes a 18.9GB Q4_0 build plus Ollama, llama.cpp and LM Studio recipes, with local contexts scaling from 4K to 256K+ depending on VRAM.
chatcodereasoningagentictool-calling
#10 · Best match
35B (3B active, MoE) · 32GB min · Q4_K_M · 21GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: InternScience Apache 2.0 agentic VLM. 35B-A3B MoE, 262K context, strong long-horizon search/tool-use benchmarks and official Q4_K_M GGUF artifacts for local workstations.
chatcodevisionagentreasoning
#11 · Best match
27B · 32GB min · Q4_K_M · 17GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Qwen 3.6 flagship dense model. Hybrid thinking mode with /think toggle for deep chain-of-thought reasoning. 128K context, 29+ languages. Significantly outperforms Qwen3.5-27B on reasoning, coding & math. Apache 2.0.
chatcodereasoningpowerquality
#12 · Best match
26B (A4B active) · 24GB min · Q4_K_M · 16GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Gemma 4 MoE flagship-for-workstations: 26B total with ~4B active parameters. 256K context and excellent quality-per-watt for local inference. Apache 2.0.
chatcodereasoningpowermultimodal
#13 · Best match
33B · 64GB min · Q4_K_M · 19.82GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official Apache 2.0 LLM-jp reasoning model with English/Japanese support, 65K GGUF context metadata and an official Q4_K_M GGUF path for local llama.cpp and LM Studio testing on 64GB+ workstations.
chatcodereasoninglong-contextmultilingual
#14 · Best match
2B · 4GB min · Q4_K_M · 1.6GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official OpenBMB compact on-device LLM with Apache 2.0 licensing, 131K context, tool-calling and coding focus, plus official Q4_K_M GGUF, Ollama, llama.cpp, Docker and OpenClaw run paths.
chatcodereasoningagentlight
#15 · Best match
8.3B (1.5B active) · 8GB min · Q4_K_M · 5.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI hybrid model built for on-device assistants. 8.3B total / 1.5B active, 128K context, tool use, GGUF, ONNX, MLX, llama.cpp and LM Studio support. Open-weight under LFM 1.0.
chatcodereasoningspeedstandard
#16 · Best match
14B · 16GB min · Q4_K_M · 9.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Established 14B Qwen 3 model for reasoning, coding and chat. Its Q4 build is a tight fit on many 16GB machines, so context length and system headroom matter.
chatcodereasoningpowergeneral
#17 · Best match
3B · 8GB min · Q4_K_M · 2.32GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.2 3B is the compact Apache 2.0 Granite reasoning model with 128K native context, thinking-mode chat, tool calling and official GGUF artifacts for laptop-class local inference.
chatcodereasoningtool-callinglight
#18 · Best match
3B multimodal · 8GB min · Q4_K_M + mmproj · 2.3GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI edge vision-language model with LFM2.5-2.6B backbone, SigLIP2 NaFlex vision encoder, 32K context, LFM 1.0 open weights and official GGUF plus llama.cpp and MLX runtime paths for local image chat and OCR.
chatvisionspeededgemultimodal