Public model card reference

ZAYA1-8B

Catalogue summary: Zyphra's Apache-2.0 reasoning MoE: 8.4B total parameters with only ~760M active, 16 experts, 131K context, Compressed Convolutional Attention and strong math/code benchmarks. Experimental for local use today: currently needs Zyphra vLLM/Transformers forks; LM Studio/GGUF/MLX support is not yet verified.

Repository editorial metadata; verify comparative claims in the linked upstream material.

Public model cardNo public GGUF verified8.4B (760M active, MoE)2026-05
Choose an app

Only verified options can be opened. Unsupported apps are clearly marked.

Compare models
LM StudioNot available for this model
UnslothNot available for this model
OllamaNot available for this model
Open on Hugging FaceModel card, licence and access details
llama.cppNot available for this model
Use with LocalClawOptional workspace after the model is installed
Source status
Public model card
Public GGUF
Not verified
Parameters
8.4B (760M active, MoE)
Access
Public card

Verified source status

LocalClaw verified a public model card for ZAYA1-8B on 2026-09-07, but no public GGUF file in that repository. The catalogue RAM and quantization fields are estimates, not a verified install path.

Only the public model card was verified; no public GGUF file was verified in that repository. Open the model card to confirm current artefacts and supported runtimes. LocalClaw does not claim a one-click LM Studio install.

chatcodereasoningmathexperimental

Source availability

01
Open the model cardUse the verified upstream card as the source of truth for this record.
02
Inspect current filesNo public GGUF file was verified in the repository during the catalogue audit.
03
Confirm a supported runtimeChoose a local runtime only after verifying the exact artefact, licence and hardware requirements upstream.

Catalogue record

  • Family: zaya
  • Parameters: 8.4B (760M active, MoE)
  • Recommended quantization: BF16 (Zyphra fork)
  • Catalogue minimum RAM: 24 GB
  • Catalogue model size: 17 GB
  • Tags: chat, code, reasoning, math, experimental

Practical limits

  • Catalogue RAM is a minimum estimate, not a guarantee for every context length or runtime.
  • Speed and memory use vary by quantization, backend, context length and system headroom.
  • Verify architecture, licence and usage restrictions in the linked upstream material before deployment.

Catalogue tags

  • chat
  • code
  • reasoning
  • math
  • experimental

Capability profile

Repository catalogue ratings used by LocalClaw's editorial rubric. They are not a standardized third-party benchmark.

speed
7
quality
8
coding
8
reasoning
9

Technical notes

Developer
Zyphra
License
Apache 2.0
Context window
131,072 tokens
Architecture
Sparse MoE with Compressed Convolutional Attention (CCA), 16 experts, top-1 MLP router and learned residual scaling

Related catalogue entries

Linked mechanically by family, shared tags and nearby RAM tier; this is not a quality ranking.

Where to go next