Mistral
How much VRAM does Mixtral 8x22B need?
Mixtral 8x22B has 141B parameters, so its weights alone take about 280 GB of VRAM at FP16 — or roughly 71 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
MoE: all 8 experts resident in memory; only 2 active per token for compute.
Total params
141B
Active params
39B
Layers
56
Hidden size
6144
Attention heads
48
KV heads (GQA)
8
Vocab size
32,000
Native context window
65,536
Native precision
BF16
Experts (total)
8
Experts active / token
2
Size Mixtral 8x22B for your workload
Opens the free calculator with Mixtral 8x22B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →