Qwen
Qwen3.6 27B
Hybrid attention: only 1 in 4 layers (Gated Attention/GQA) carries a KV cache that grows with context; the rest (Gated DeltaNet) hold a small fixed-size recurrent state instead, not modeled here. Native context extends to ~1.01M via YaRN in supported serving stacks.
Total params
27B
Active params
27B
Layers
64
Hidden size
5120
Attention heads
24
KV heads (GQA)
4
Vocab size
248,320
Native context window
262,144
Native precision
BF16
Layers with growing KV cache
25%
Size this model
Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.
Size this model →