#1 · Best match
8.8B · 8GB min · Q4_K_M · 5.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.2 8B instruct model with Apache 2.0 weights, 128K context, thinking-mode chat template, tool calling and practical GGUF plus MLX paths for everyday local machines.
chatcodereasoningtool-callingstandard
#2 · Best match
9B · 16GB min · Q4_K_M · 5.63GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official MIT reasoning model from Ornith AI with a 262K native context window, tool-calling focus, and official Q4_K_M GGUF, MLX and Ollama/llama.cpp paths for local coding-agent experiments on 16GB+ machines.
chatcodereasoningagenticlong-context
#3 · Best match
2B · 4GB min · Q4_K_M · 1.6GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official OpenBMB compact on-device LLM with Apache 2.0 licensing, 131K context, tool-calling and coding focus, plus official Q4_K_M GGUF, Ollama, llama.cpp, Docker and OpenClaw run paths.
chatcodereasoningagentlight
#4 · Best match
8.3B (1.5B active) · 8GB min · Q4_K_M · 5.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI hybrid model built for on-device assistants. 8.3B total / 1.5B active, 128K context, tool use, GGUF, ONNX, MLX, llama.cpp and LM Studio support. Open-weight under LFM 1.0.
chatcodereasoningspeedstandard
#5 · Best match
14B · 16GB min · Q4_K_M · 9.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Established 14B Qwen 3 model for reasoning, coding and chat. Its Q4 build is a tight fit on many 16GB machines, so context length and system headroom matter.
chatcodereasoningpowergeneral
#6 · Best match
3B · 8GB min · Q4_K_M · 2.32GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.2 3B is the compact Apache 2.0 Granite reasoning model with 128K native context, thinking-mode chat, tool calling and official GGUF artifacts for laptop-class local inference.
chatcodereasoningtool-callinglight
#7 · Best match
3B multimodal · 8GB min · Q4_K_M + mmproj · 2.3GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI edge vision-language model with LFM2.5-2.6B backbone, SigLIP2 NaFlex vision encoder, 32K context, LFM 1.0 open weights and official GGUF plus llama.cpp and MLX runtime paths for local image chat and OCR.
chatvisionspeededgemultimodal
#8 · Best match
7.9B (1.3B active, MoE) · 8GB min · Q4_K_M · 4.82GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official InclusionAI MIT hybrid-reasoning MoE with 131K context, 1.3B active parameters and a validated GGUF path through stock llama.cpp builds from the 2026-08-17 bailingmoe3 merge.
chatcodereasoningagenticlong-context
#9 · Best match
2.7B · 8GB min · Q4_K_M · 1.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI compact hybrid model with 128K context, LFM 1.0 open weights, official GGUF, ONNX and MLX artifacts, and practical llama.cpp / LM Studio paths for 8GB-class local machines.
chatcodereasoninglightspeed
#10 · Best match
8B · 8GB min · Q4_K_M · 5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.1 long-context instruct model. Apache 2.0, 131K context, tool calling, RAG, code tasks, multilingual dialog and business assistant workflows on normal 8-16 GB machines.
chatcodereasoningstandardgeneral
#11 · Best match
4B · 16GB min · BF16 GGUF · 7.67GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Spark-X2.5-4B is an Apache 2.0 compact general-purpose model from XHToken with a hybrid attention architecture, 1M-token native context, multilingual coverage and official GGUF artifacts for local llama.cpp, Ollama and LM Studio-compatible workflows.
chatcodereasoningtool-callinglong-context
#12 · Best match
12B · 16GB min · Q4_K_M · 8.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Google DeepMind 12B unified multimodal model. Text, image, audio and video inputs, 256K context, Apache 2.0, and a strong local sweet spot for 16-32 GB machines.
chatvisionaudiocodereasoning
#13 · Best match
7B · 16GB min · Q4_K_M · 4.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Ai2 fully open 7B OLMo 3 instruct model with Apache 2.0 licensing, transparent training artifacts and established GGUF options from Unsloth, bartowski and LM Studio for everyday local machines.
chatcodereasoningstandardopen-data
#14 · Best match
8B · 8GB min · Q5_K_M · 5.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Established Qwen 3 8B model with thinking mode, strong general capability and broad GGUF runtime support.
chatcodestandardgeneralreasoning
#15 · Best match
14B · 16GB min · Q4_K_M · 8.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Mistral AI larger Ministral 3 instruct model. Apache 2.0, official GGUF availability, better quality ceiling than the 3B/8B variants while staying practical on 16-32GB workstations.
chatvisionpowerreasoningmultilingual
#16 · Best match
8B · 12GB min · Q4_K_M · 5.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Qwen 3 vision-language model. Strong OCR, document understanding, chart & UI reasoning. 128K context with native image+video inputs. Apache 2.0.
visionchatmultimodalstandard
#17 · Best match
14B · 16GB min · Q6_K · 12GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Microsoft's full Phi-4. Compact powerhouse with exceptional reasoning and coding for its size. MIT licensed.
chatcodepowerreasoning
#18 · Best match
12B · 12GB min · Q5_K_M · 7.1GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Mistral x NVIDIA 128K context model. Excellent for long documents and conversations. 2.7M downloads.
chatgeneralstandard