Qwen
Qwen3 32B
Native context is 40,960 tokens; extendable to 131,072 via YaRN rope scaling in supported serving stacks.
Total params
32.8B
Active params
32.8B
Layers
64
Hidden size
5120
Attention heads
64
KV heads (GQA)
8
Vocab size
151,936
Native context window
40,960
Native precision
BF16
Size this model
Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.
Size this model →