Qwen
How much VRAM does Qwen3 235B-A22B need?
Qwen3 235B-A22B has 235B parameters, so its weights alone take about 470 GB of VRAM at FP16 — or roughly 120 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
MoE: all 128 experts resident in memory; 8 active per token for compute.
Total params
235B
Active params
22B
Layers
94
Hidden size
4096
Attention heads
64
KV heads (GQA)
4
Vocab size
151,936
Native context window
40,960
Native precision
BF16
Experts (total)
128
Experts active / token
8
Size Qwen3 235B-A22B for your workload
Opens the free calculator with Qwen3 235B-A22B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →