#1 · Best match
8.8B · 8GB min · Q4_K_M · 5.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.2 8B instruct model with Apache 2.0 weights, 128K context, thinking-mode chat template, tool calling and practical GGUF plus MLX paths for everyday local machines.
chatcodereasoningtool-callingstandard
#2 · Best match
9B · 16GB min · Q4_K_M · 5.63GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official MIT reasoning model from Ornith AI with a 262K native context window, tool-calling focus, and official Q4_K_M GGUF, MLX and Ollama/llama.cpp paths for local coding-agent experiments on 16GB+ machines.
chatcodereasoningagenticlong-context
#3 · Best match
2B · 4GB min · Q4_K_M · 1.6GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official OpenBMB compact on-device LLM with Apache 2.0 licensing, 131K context, tool-calling and coding focus, plus official Q4_K_M GGUF, Ollama, llama.cpp, Docker and OpenClaw run paths.
chatcodereasoningagentlight
#4 · Best match
8.3B (1.5B active) · 8GB min · Q4_K_M · 5.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI hybrid model built for on-device assistants. 8.3B total / 1.5B active, 128K context, tool use, GGUF, ONNX, MLX, llama.cpp and LM Studio support. Open-weight under LFM 1.0.
chatcodereasoningspeedstandard
#5 · Best match
3B · 8GB min · Q4_K_M · 2.32GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.2 3B is the compact Apache 2.0 Granite reasoning model with 128K native context, thinking-mode chat, tool calling and official GGUF artifacts for laptop-class local inference.
chatcodereasoningtool-callinglight
#6 · Best match
3B multimodal · 8GB min · Q4_K_M + mmproj · 2.3GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI edge vision-language model with LFM2.5-2.6B backbone, SigLIP2 NaFlex vision encoder, 32K context, LFM 1.0 open weights and official GGUF plus llama.cpp and MLX runtime paths for local image chat and OCR.
chatvisionspeededgemultimodal
#7 · Best match
7.9B (1.3B active, MoE) · 8GB min · Q4_K_M · 4.82GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official InclusionAI MIT hybrid-reasoning MoE with 131K context, 1.3B active parameters and a validated GGUF path through stock llama.cpp builds from the 2026-08-17 bailingmoe3 merge.
chatcodereasoningagenticlong-context
#8 · Best match
2.7B · 8GB min · Q4_K_M · 1.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI compact hybrid model with 128K context, LFM 1.0 open weights, official GGUF, ONNX and MLX artifacts, and practical llama.cpp / LM Studio paths for 8GB-class local machines.
chatcodereasoninglightspeed
#9 · Best match
8B · 8GB min · Q4_K_M · 5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.1 long-context instruct model. Apache 2.0, 131K context, tool calling, RAG, code tasks, multilingual dialog and business assistant workflows on normal 8-16 GB machines.
chatcodereasoningstandardgeneral
#10 · Best match
7B · 16GB min · Q4_K_M · 4.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Ai2 fully open 7B OLMo 3 instruct model with Apache 2.0 licensing, transparent training artifacts and established GGUF options from Unsloth, bartowski and LM Studio for everyday local machines.
chatcodereasoningstandardopen-data
#11 · Best match
8B · 8GB min · Q5_K_M · 5.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Established Qwen 3 8B model with thinking mode, strong general capability and broad GGUF runtime support.
chatcodestandardgeneralreasoning
#12 · Best match
8B · 12GB min · Q4_K_M · 5.2GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Qwen 3 vision-language model. Strong OCR, document understanding, chart & UI reasoning. 128K context with native image+video inputs. Apache 2.0.
visionchatmultimodalstandard
#13 · Best match
12B · 12GB min · Q5_K_M · 7.1GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Mistral x NVIDIA 128K context model. Excellent for long documents and conversations. 2.7M downloads.
chatgeneralstandard
#14 · Best match
1.7B · 8GB min · BF16 GGUF · 3.19GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Spark-X2.5-1.7B is the smaller Apache 2.0 Spark-X2.5 release, tuned for lightweight conversation, coding, reasoning and agentic workflows with a 1M-token native context claim and official GGUF local runtime artifacts.
chatcodereasoningtool-callinglong-context
#15 · Best match
1.8B · 4GB min · Q8_0 · 1.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: DFM-Mimir is a Danish and English 1B-class HRM language model trained from scratch by Danish Foundation Models, with Apache 2.0 weights, permissible-data positioning and GGUF artifacts for lightweight local inference.
chatreasoninglightmultilingualgeneral
#16 · Best match
4B · 8GB min · Q4_K_M · 2.71GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: InternScience compact dense agent model with Apache 2.0 licensing, 262K context and official Q4_K_M GGUF artifacts for 8GB-class local assistants.
chatcodereasoningvisionagentic
#17 · Best match
3B · 4GB min · Q4_K_M · 2.1GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.1 compact long-context instruct model. Apache 2.0, 131K context, tool calling, RAG and code tasks, with an official Q4_K_M GGUF for practical 4-8 GB local machines.
chatcodereasoninglightspeed
#18 · Best match
9B · 8GB min · Q4_K_M · 5.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Compact Ornith 1.0 GGUF variant from DeepReinforce for agentic coding experiments on consumer hardware. MIT licensed and much more practical than the frontier 397B release.
chatcodereasoningspeedagentic