Qwen
How much VRAM does Qwen3 30B-A3B need?
Qwen3 30B-A3B has 30.5B parameters, so its weights alone take about 61 GB of VRAM at FP16 — or roughly 15 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
MoE: all 128 experts resident in memory; 8 active per token for compute.
Total params
30.5B
Active params
3.3B
Layers
48
Hidden size
2048
Attention heads
32
KV heads (GQA)
4
Vocab size
151,936
Native context window
40,960
Native precision
BF16
Experts (total)
128
Experts active / token
8
Size Qwen3 30B-A3B for your workload
Opens the free calculator with Qwen3 30B-A3B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →