Llama

How much VRAM does Llama 4 Scout (17B active / 109B total) need?

Llama 4 Scout (17B active / 109B total) has 109B parameters, so its weights alone take about 220 GB of VRAM at FP16 — or roughly 55 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.

MoE: all 16 experts resident in memory; 1 active per token for compute. 10M-token context window is exceptional — most serving stacks size for a much smaller practical window.

Total params
109B
Active params
17B
Layers
48
Hidden size
5120
Attention heads
40
KV heads (GQA)
8
Vocab size
202,048
Native context window
10,485,760
Native precision
BF16
Experts (total)
16
Experts active / token
1
Size Llama 4 Scout (17B active / 109B total) for your workload

Opens the free calculator with Llama 4 Scout (17B active / 109B total) loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.

Size this model →