Gated model card reference

Llama 4 Maverick (17B/400B MoE)

Catalogue summary: Meta Llama 4 Maverick — 128-expert MoE flagship. Matches or beats GPT-4o and Gemini 2.0 Flash on reasoning, coding and multimodal benchmarks. 1M-token context. Server-grade hardware only. Llama 4 Community License.

Repository editorial metadata; verify comparative claims in the linked upstream material.

Gated model cardNo public GGUF verified400B (17B active, 128 experts)2025-04
Choose an app

Only verified options can be opened. Unsupported apps are clearly marked.

Compare models
LM StudioNot available for this model
UnslothNot available for this model
OllamaNot available for this model
Open on Hugging FaceModel card, licence and access details
llama.cppNot available for this model
Use with LocalClawOptional workspace after the model is installed
Source status
Gated model card
Public GGUF
Not verified
Parameters
400B (17B active, 128 experts)
Access
Approval required

Verified source status

LocalClaw verified a gated model card for Llama 4 Maverick (17B/400B MoE) on 2026-09-07. Access approval or licence acceptance is required, and no public GGUF install path is claimed.

The model card is gated and may require account approval or licence acceptance. No public GGUF file was verified, so LocalClaw does not publish an LM Studio or one-click installation path.

chatvisionreasoningmultimodalquality

Access and source status

01
Open the model cardUse the verified upstream card as the source of truth for this record.
02
Request accessAccount approval or licence acceptance may be required before the files can be inspected.
03
Confirm a supported runtimeChoose a local runtime only after verifying the exact artefact, licence and hardware requirements upstream.

Catalogue record

  • Family: llama
  • Parameters: 400B (17B active, 128 experts)
  • Recommended quantization: Q4_K_M
  • Catalogue minimum RAM: 384 GB
  • Catalogue model size: 240 GB
  • Tags: chat, vision, reasoning, multimodal, quality

Practical limits

  • Catalogue RAM is a minimum estimate, not a guarantee for every context length or runtime.
  • Speed and memory use vary by quantization, backend, context length and system headroom.
  • Verify architecture, licence and usage restrictions in the linked upstream material before deployment.

Catalogue tags

  • chat
  • vision
  • reasoning
  • multimodal
  • quality

Capability profile

Repository catalogue ratings used by LocalClaw's editorial rubric. They are not a standardized third-party benchmark.

speed
2
quality
10
coding
10
reasoning
10

Technical notes

Developer
See upstream repository
License
See upstream repository
Context window
See upstream repository
Architecture
See upstream repository

Related catalogue entries

Linked mechanically by family, shared tags and nearby RAM tier; this is not a quality ranking.

Where to go next