#1Best match
Meta Superintelligence Lab local agent model with text+image input, 131K context, Apache 2.0 weights and official GGUF/ExecuTorch artifacts. The K-Quant 17GB build targets 24GB machines; 32GB is safer for vision and long-context sessions.
Parameters29.8B multimodalMinimum RAM24GBQuantizationK-Quant 17GB Q4_K_MModel size17GB
View model details →#2Best match
Official Qwen dense 27B vision-language release with Apache 2.0 weights, 262K native context, thinking controls and strong agentic coding benchmarks. Practical local path through Unsloth and LM Studio-compatible GGUF artifacts.
Parameters27BMinimum RAM32GBQuantizationQ4_K_MModel size16.8GB
View model details →#3Best match
IBM Granite 4.2 8B instruct model with Apache 2.0 weights, 128K context, thinking-mode chat template, tool calling and practical GGUF plus MLX paths for everyday local machines.
Parameters8.8BMinimum RAM8GBQuantizationQ4_K_MModel size5.2GB
View model details →#4Best match
Official MIT reasoning model from Ornith AI with a 262K native context window, tool-calling focus, and official Q4_K_M GGUF, MLX and Ollama/llama.cpp paths for local coding-agent experiments on 16GB+ machines.
Parameters9BMinimum RAM16GBQuantizationQ4_K_MModel size5.63GB
View model details →#5Best match
Qwen 3.6 flagship dense model. Hybrid thinking mode with /think toggle for deep chain-of-thought reasoning. 128K context, 29+ languages. Significantly outperforms Qwen3.5-27B on reasoning, coding & math. Apache 2.0.
Parameters27BMinimum RAM32GBQuantizationQ4_K_MModel size17GB
View model details →#6Best match
Gemma 4 MoE flagship-for-workstations: 26B total with ~4B active parameters. 256K context and excellent quality-per-watt for local inference. Apache 2.0.
Parameters26B (A4B active)Minimum RAM24GBQuantizationQ4_K_MModel size16GB
View model details →#7Best match
Official OpenBMB compact on-device LLM with Apache 2.0 licensing, 131K context, tool-calling and coding focus, plus official Q4_K_M GGUF, Ollama, llama.cpp, Docker and OpenClaw run paths.
Parameters2BMinimum RAM4GBQuantizationQ4_K_MModel size1.6GB
View model details →#8Best match
Liquid AI hybrid model built for on-device assistants. 8.3B total / 1.5B active, 128K context, tool use, GGUF, ONNX, MLX, llama.cpp and LM Studio support. Open-weight under LFM 1.0.
Parameters8.3B (1.5B active)Minimum RAM8GBQuantizationQ4_K_MModel size5.2GB
View model details →#9Best match
Established 14B Qwen 3 model for reasoning, coding and chat. Its Q4 build is a tight fit on many 16GB machines, so context length and system headroom matter.
Parameters14BMinimum RAM16GBQuantizationQ4_K_MModel size9.5GB
View model details →