Qwen
How much VRAM does Qwen3 32B need?
Qwen3 32B has 32.8B parameters, so its weights alone take about 66 GB of VRAM at FP16 — or roughly 16 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
Native context is 40,960 tokens; extendable to 131,072 via YaRN rope scaling in supported serving stacks.
Total params
32.8B
Active params
32.8B
Layers
64
Hidden size
5120
Attention heads
64
KV heads (GQA)
8
Vocab size
151,936
Native context window
40,960
Native precision
BF16
Size Qwen3 32B for your workload
Opens the free calculator with Qwen3 32B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →