Qwen
Qwen3 30B-A3B
MoE: all 128 experts resident in memory; 8 active per token for compute.
Total params
30.5B
Active params
3.3B
Layers
48
Hidden size
2048
Attention heads
32
KV heads (GQA)
4
Vocab size
151,936
Native context window
40,960
Native precision
BF16
Experts (total)
128
Experts active / token
8
Size this model
Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.
Size this model →