Qwen

Qwen3.6 35B-A3B

MoE: 256 experts resident in memory, 8 routed + 1 shared active per token for compute. Hybrid attention: only 1 in 4 layers (Gated Attention/GQA) carries a growing KV cache; the rest (Gated DeltaNet) hold a small fixed-size recurrent state instead, not modeled here.

Total params
35B
Active params
3B
Layers
40
Hidden size
2048
Attention heads
16
KV heads (GQA)
2
Vocab size
248,320
Native context window
262,144
Native precision
BF16
Experts (total)
256
Experts active / token
9
Layers with growing KV cache
25%
Size this model

Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.

Size this model →