Qwen

How much VRAM does Qwen3.6 35B-A3B need?

Qwen3.6 35B-A3B has 35B parameters, so its weights alone take about 70 GB of VRAM at FP16 — or roughly 18 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.

MoE: 256 experts resident in memory, 8 routed + 1 shared active per token for compute. Hybrid attention: only 1 in 4 layers (Gated Attention/GQA) carries a growing KV cache; the rest (Gated DeltaNet) hold a small fixed-size recurrent state instead, not modeled here.

Total params
35B
Active params
3B
Layers
40
Hidden size
2048
Attention heads
16
KV heads (GQA)
2
Vocab size
248,320
Native context window
262,144
Native precision
BF16
Experts (total)
256
Experts active / token
9
Layers with growing KV cache
25%
Size Qwen3.6 35B-A3B for your workload

Opens the free calculator with Qwen3.6 35B-A3B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.

Size this model →