Mistral
How much VRAM does Mixtral 8x7B need?
Mixtral 8x7B has 46.7B parameters, so its weights alone take about 93 GB of VRAM at FP16 — or roughly 23 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
MoE: all 8 experts resident in memory; only 2 active per token for compute.
Total params
46.7B
Active params
12.9B
Layers
32
Hidden size
4096
Attention heads
32
KV heads (GQA)
8
Vocab size
32,000
Native context window
32,768
Native precision
BF16
Experts (total)
8
Experts active / token
2
Size Mixtral 8x7B for your workload
Opens the free calculator with Mixtral 8x7B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →