#1Best match
Official MIT 35B MoE reasoning model from Ornith AI with about 3B active parameters, 262K context, strong agentic-coding positioning and official Q4_K_M GGUF plus MLX/Ollama/llama.cpp local paths for larger workstations.
Parameters35B (3B active, MoE)Minimum RAM48GBQuantizationQ4_K_MModel size21.72GB
View model details →#2Best match
Meta Superintelligence Lab local agent model with text+image input, 131K context, Apache 2.0 weights and official GGUF/ExecuTorch artifacts. The K-Quant 17GB build targets 24GB machines; 32GB is safer for vision and long-context sessions.
Parameters29.8B multimodalMinimum RAM24GBQuantizationK-Quant 17GB Q4_K_MModel size17GB
View model details →#3Best match
Official Qwen dense 27B vision-language release with Apache 2.0 weights, 262K native context, thinking controls and strong agentic coding benchmarks. Practical local path through Unsloth and LM Studio-compatible GGUF artifacts.
Parameters27BMinimum RAM32GBQuantizationQ4_K_MModel size16.8GB
View model details →#4Best match
IBM Granite 4.2 30B brings the permissive Apache 2.0 Granite stack to workstation-class local reasoning, RAG, coding and tool-use workflows with GGUF and MLX community artifacts.
Parameters29.3BMinimum RAM32GBQuantizationQ4_K_MModel size18GB
View model details →#5Best match
InclusionAI's MIT-licensed instruct MoE optimized for fast agent workloads. 104B total parameters, only 7.4B active, hybrid linear attention, 262K context and strong tool-use / multi-step execution with high token efficiency.
Parameters104B (7.4B active)Minimum RAM80GBQuantizationQ4_K_MModel size65GB
View model details →#6Best match
IBM Granite 4.2 8B instruct model with Apache 2.0 weights, 128K context, thinking-mode chat template, tool calling and practical GGUF plus MLX paths for everyday local machines.
Parameters8.8BMinimum RAM8GBQuantizationQ4_K_MModel size5.2GB
View model details →#7Best match
Qwen Team open-weight MoE for agentic coding and multimodal work. 35B total / 3B active, 262K native context, Apache 2.0, and strong GGUF availability through Unsloth and LM Studio-compatible artifacts.
Parameters35B (3B active, MoE)Minimum RAM32GBQuantizationQ4_K_MModel size19GB
View model details →#8Best match
Large MoE model with only 10B active params. 60% cheaper to run than Qwen3-Max. 256K context. Top-tier reasoning, coding and multilingual. Hybrid think/non-think. Apache 2.0.
Parameters122B (10B active)Minimum RAM80GBQuantizationQ4_K_MModel size65GB
View model details →#9Best match
Qwen 3 dense 32B open-weight model with hybrid thinking and non-thinking modes, strong reasoning and coding support, and a practical Q4_K_M GGUF path for 32GB-class local machines.
Parameters32BMinimum RAM32GBQuantizationQ4_K_MModel size20GB
View model details →