#1Best match
Official MIT DeepSeek V4 Flash multimodal experiment with image understanding, 1M context and Unsloth Dynamic GGUF artifacts. The lightest practical GGUF is roughly 82-97GB, while higher-quality Q4/Q8 builds are about 155-162GB, so this belongs on large-memory workstations.
Parameters284B (13B active, multimodal MoE)Minimum RAM128GBQuantizationUD-Q2_K_XLModel size97GB
View model details →#2Best match
Official MIT DeepSeek V4 Flash successor release with stronger agentic coding, DSpark speculative decoding support and a practical Unsloth Dynamic GGUF path. Still a large workstation/server local model: Q4 is about 155GB and Q8 is about 162GB.
Parameters284B (13B active)Minimum RAM256GBQuantizationUD-Q4_K_XLModel size155GB
View model details →#3Best match
Official MIT 35B MoE reasoning model from Ornith AI with about 3B active parameters, 262K context, strong agentic-coding positioning and official Q4_K_M GGUF plus MLX/Ollama/llama.cpp local paths for larger workstations.
Parameters35B (3B active, MoE)Minimum RAM48GBQuantizationQ4_K_MModel size21.72GB
View model details →#4Best match
Meta Superintelligence Lab local agent model with text+image input, 131K context, Apache 2.0 weights and official GGUF/ExecuTorch artifacts. The K-Quant 17GB build targets 24GB machines; 32GB is safer for vision and long-context sessions.
Parameters29.8B multimodalMinimum RAM24GBQuantizationK-Quant 17GB Q4_K_MModel size17GB
View model details →#5Best match
Official Qwen dense 27B vision-language release with Apache 2.0 weights, 262K native context, thinking controls and strong agentic coding benchmarks. Practical local path through Unsloth and LM Studio-compatible GGUF artifacts.
Parameters27BMinimum RAM32GBQuantizationQ4_K_MModel size16.8GB
View model details →#6Best match
IBM Granite 4.2 30B brings the permissive Apache 2.0 Granite stack to workstation-class local reasoning, RAG, coding and tool-use workflows with GGUF and MLX community artifacts.
Parameters29.3BMinimum RAM32GBQuantizationQ4_K_MModel size18GB
View model details →#7Best match
InclusionAI's MIT-licensed instruct MoE optimized for fast agent workloads. 104B total parameters, only 7.4B active, hybrid linear attention, 262K context and strong tool-use / multi-step execution with high token efficiency.
Parameters104B (7.4B active)Minimum RAM80GBQuantizationQ4_K_MModel size65GB
View model details →#8Best match
IBM Granite 4.2 8B instruct model with Apache 2.0 weights, 128K context, thinking-mode chat template, tool calling and practical GGUF plus MLX paths for everyday local machines.
Parameters8.8BMinimum RAM8GBQuantizationQ4_K_MModel size5.2GB
View model details →#9Best match
Z.ai flagship open model for long-horizon coding, reasoning and agentic work. 744B total, 40B active, 1M-token context, MIT license. Unsloth Dynamic GGUF makes it technically local, but it needs workstation/server-class memory: ~245GB total memory for 2-bit and 372GB+ for 4-bit.
Parameters744B (40B active)Minimum RAM256GBQuantizationUD-IQ2_MModel size239GB
View model details →