#1 · Best match
2B · 4GB min · Q4_K_M · 1.6GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Official OpenBMB compact on-device LLM with Apache 2.0 licensing, 131K context, tool-calling and coding focus, plus official Q4_K_M GGUF, Ollama, llama.cpp, Docker and OpenClaw run paths.
chatcodereasoningagentlight
#2 · Best match
3B · 8GB min · Q4_K_M · 2.32GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.2 3B is the compact Apache 2.0 Granite reasoning model with 128K native context, thinking-mode chat, tool calling and official GGUF artifacts for laptop-class local inference.
chatcodereasoningtool-callinglight
#3 · Best match
3B multimodal · 8GB min · Q4_K_M + mmproj · 2.3GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI edge vision-language model with LFM2.5-2.6B backbone, SigLIP2 NaFlex vision encoder, 32K context, LFM 1.0 open weights and official GGUF plus llama.cpp and MLX runtime paths for local image chat and OCR.
chatvisionspeededgemultimodal
#4 · Best match
2.7B · 8GB min · Q4_K_M · 1.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Liquid AI compact hybrid model with 128K context, LFM 1.0 open weights, official GGUF, ONNX and MLX artifacts, and practical llama.cpp / LM Studio paths for 8GB-class local machines.
chatcodereasoninglightspeed
#5 · Best match
1.8B · 4GB min · Q8_0 · 1.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: DFM-Mimir is a Danish and English 1B-class HRM language model trained from scratch by Danish Foundation Models, with Apache 2.0 weights, permissible-data positioning and GGUF artifacts for lightweight local inference.
chatreasoninglightmultilingualgeneral
#6 · Best match
4B · 8GB min · Q4_K_M · 2.71GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: InternScience compact dense agent model with Apache 2.0 licensing, 262K context and official Q4_K_M GGUF artifacts for 8GB-class local assistants.
chatcodereasoningvisionagentic
#7 · Best match
3B · 4GB min · Q4_K_M · 2.1GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM Granite 4.1 compact long-context instruct model. Apache 2.0, 131K context, tool calling, RAG and code tasks, with an official Q4_K_M GGUF for practical 4-8 GB local machines.
chatcodereasoninglightspeed
#8 · Best match
4B · 6GB min · Q5_K_M · 2.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: NVIDIA fine-tune of Llama 3.1 with hybrid /think and /no_think modes, 128K context and Apache 2.0 licensing. A compact local option for reasoning and chat.
chatlightspeedreasoning
#9 · Best match
4B · 6GB min · Q5_K_M · 2.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: NVIDIA compact hybrid model distilled from a 9B teacher, with hybrid attention and SSM layers. A lightweight option for local chat and reasoning under the NVIDIA Open Model License.
chatlightspeedreasoning
#10 · Best match
1B · 4GB min · Q4_K_M · 0.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: OpenBMB compact on-device LLM with Apache 2.0 licensing, 128K context, tool-calling focus and official GGUF plus MLX artifacts for laptops and edge devices.
chatcodereasoninglightspeed
#11 · Best match
1.5B · 4GB min · Q4_K_M · 0.9GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IbnSina-1.5B is a Persian-first 1.48B Llama-compatible language model trained from scratch on a Persian-heavy corpus, with Apache 2.0 weights and GGUF artifacts for laptop, phone, Ollama, LM Studio and llama.cpp use.
chatlightspeedmultilingualgeneral
#12 · Best match
2B · 4GB min · Q5_K_M · 1.4GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: IBM ultra-efficient 2B. Best-in-class among small models for tool calling & structured output. Perfect for on-device RAG and agents. 128K context. Apache 2.0.
chatlightedgespeedcode
#13 · Best match
4B · 6GB min · Q4_K_M · 3GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Compact Qwen 3.5 model with hybrid thinking, 256K context and multilingual support. Runs in an 8GB-class local setup with the listed quantization. Apache 2.0.
chatcodereasoningspeedgeneral
#14 · Best match
3B · 4GB min · Q4_K_M · 2.1GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Mistral AI compact multimodal instruct model. Apache 2.0, strong local app support through official GGUF, LM Studio, Ollama and llama.cpp artifacts. Practical on normal laptops.
chatvisionlightspeedgeneral
#15 · Best match
4B · 6GB min · Q5_K_M · 2.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Google Gemma for phones/tablets/laptops. Optimized for mobile and edge. 552K downloads.
chatlightspeed
#16 · Best match
4B · 4GB min · Q5_K_M · 2.8GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Alibaba's think-then-answer model. Built-in chain-of-thought reasoning at just 4B params.
chatcodelightspeedreasoning
#17 · Best match
4B · 8GB min · Q5_K_M · 3GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Google's multimodal gem. Understands text AND images natively. Great quality-to-size ratio.
chatvisionstandardgeneral
#18 · Best match
3.8B · 4GB min · Q5_K_M · 2.5GB
Why it fits: Comfortable memory headroom. General match. Catalogue summary: Microsoft's latest small miracle. Punches way above its weight in reasoning & code.
chatcodelightspeed