Open-weight MoE

Nemotron 3.5 Lightning 30B-A3B

Catalogue summary: NVIDIA OpenMDW-1.1 hybrid Mamba/MoE/attention model for local agentic inference. The official GGUF path from ggml-org includes a 18.9GB Q4_0 build plus Ollama, llama.cpp and LM Studio recipes, with local contexts scaling from 4K to 256K+ depending on VRAM.

Repository editorial metadata; verify comparative claims in the linked upstream material.

48 GB catalogue minimum48 GB RAMQ4_0Coding assistant
Parameters
30B (3B active, MoE)
Minimum RAM
48 GB
Model size
18.9 GB
Quantization
Q4_0

Can Nemotron 3.5 Lightning 30B-A3B run locally?

Nemotron 3.5 Lightning 30B-A3B has a catalogue minimum of 48 GB RAM with Q4_0. Actual memory use and speed vary by context length, runtime, backend and system headroom.

Use nvidia-nemotron-3.5-lightning-30b-a3b as the catalogue search term in a compatible runtime, and confirm the available format on the upstream repository before download.

chatcodereasoningagentictool-callinglong-contextpowergeneral

Install path

01
Check RAM fitMinimum 48 GB RAM. Start with the Q4_0 quant.
02
Load the modelSearch nvidia-nemotron-3.5-lightning-30b-a3b in LM Studio.
03
Control locallyUse LocalClaw to manage models, agents, chat, channels and scheduled OpenClaw work.

Catalogue record

  • Family: nemotron
  • Parameters: 30B (3B active, MoE)
  • Recommended quantization: Q4_0
  • Catalogue minimum RAM: 48 GB
  • Catalogue model size: 18.9 GB
  • Tags: chat, code, reasoning, agentic, tool-calling, long-context, power, general

Practical limits

  • Catalogue RAM is a minimum estimate, not a guarantee for every context length or runtime.
  • Speed and memory use vary by quantization, backend, context length and system headroom.
  • Verify architecture, licence and usage restrictions in the linked upstream material before deployment.

Catalogue tags

  • chat
  • code
  • reasoning
  • agentic
  • tool-calling
  • long-context

Capability profile

Repository catalogue ratings used by LocalClaw's editorial rubric. They are not a standardized third-party benchmark.

speed
7
quality
8
coding
8
reasoning
8

Technical notes

Developer
NVIDIA
License
OpenMDW-1.1
Context window
1,048,576 tokens
Architecture
Hybrid 30B mixture-of-experts model with about 3B active parameters per token, interleaved Mamba-2 layers, MoE layers, selected attention layers and Multi-Token Prediction support. NVIDIA publishes BF16 and NVFP4 checkpoints, while ggml-org provides the official GGUF path for llama.cpp-compatible local runtimes.

This model fits these next steps

Hardware fit is based on LocalClaw's RAM tier, model size and quantization metadata. Always leave memory headroom for your OS and runtime.

Related catalogue entries

Linked mechanically by family, shared tags and nearby RAM tier; this is not a quality ranking.

Where to go next