Llama
Llama 4 Scout (17B active / 109B total)
MoE: all 16 experts resident in memory; 1 active per token for compute. 10M-token context window is exceptional — most serving stacks size for a much smaller practical window.
Total params
109B
Active params
17B
Layers
48
Hidden size
5120
Attention heads
40
KV heads (GQA)
8
Vocab size
202,048
Native context window
10,485,760
Native precision
BF16
Experts (total)
16
Experts active / token
1
Size this model
Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.
Size this model →