#1Best match
IBM Granite 4.2 8B instruct model with Apache 2.0 weights, 128K context, thinking-mode chat template, tool calling and practical GGUF plus MLX paths for everyday local machines.
Parameters8.8BMinimum RAM8GBQuantizationQ4_K_MModel size5.2GB
View model details →#2Best match
Official MIT reasoning model from Ornith AI with a 262K native context window, tool-calling focus, and official Q4_K_M GGUF, MLX and Ollama/llama.cpp paths for local coding-agent experiments on 16GB+ machines.
Parameters9BMinimum RAM16GBQuantizationQ4_K_MModel size5.63GB
View model details →#3Best match
Official OpenBMB compact on-device LLM with Apache 2.0 licensing, 131K context, tool-calling and coding focus, plus official Q4_K_M GGUF, Ollama, llama.cpp, Docker and OpenClaw run paths.
Parameters2BMinimum RAM4GBQuantizationQ4_K_MModel size1.6GB
View model details →#4Best match
Liquid AI hybrid model built for on-device assistants. 8.3B total / 1.5B active, 128K context, tool use, GGUF, ONNX, MLX, llama.cpp and LM Studio support. Open-weight under LFM 1.0.
Parameters8.3B (1.5B active)Minimum RAM8GBQuantizationQ4_K_MModel size5.2GB
View model details →#5Best match
IBM Granite 4.2 3B is the compact Apache 2.0 Granite reasoning model with 128K native context, thinking-mode chat, tool calling and official GGUF artifacts for laptop-class local inference.
Parameters3BMinimum RAM8GBQuantizationQ4_K_MModel size2.32GB
View model details →#6Best match
Liquid AI edge vision-language model with LFM2.5-2.6B backbone, SigLIP2 NaFlex vision encoder, 32K context, LFM 1.0 open weights and official GGUF plus llama.cpp and MLX runtime paths for local image chat and OCR.
Parameters3B multimodalMinimum RAM8GBQuantizationQ4_K_M + mmprojModel size2.3GB
View model details →#7Best match
Official InclusionAI MIT hybrid-reasoning MoE with 131K context, 1.3B active parameters and a validated GGUF path through stock llama.cpp builds from the 2026-08-17 bailingmoe3 merge.
Parameters7.9B (1.3B active, MoE)Minimum RAM8GBQuantizationQ4_K_MModel size4.82GB
View model details →#8Best match
Liquid AI compact hybrid model with 128K context, LFM 1.0 open weights, official GGUF, ONNX and MLX artifacts, and practical llama.cpp / LM Studio paths for 8GB-class local machines.
Parameters2.7BMinimum RAM8GBQuantizationQ4_K_MModel size1.8GB
View model details →#9Best match
IBM Granite 4.1 long-context instruct model. Apache 2.0, 131K context, tool calling, RAG, code tasks, multilingual dialog and business assistant workflows on normal 8-16 GB machines.
Parameters8BMinimum RAM8GBQuantizationQ4_K_MModel size5GB
View model details →