Mistral
Mixtral 8x7B
MoE: all 8 experts resident in memory; only 2 active per token for compute.
Total params
46.7B
Active params
12.9B
Layers
32
Hidden size
4096
Attention heads
32
KV heads (GQA)
8
Vocab size
32,000
Native context window
32,768
Native precision
BF16
Experts (total)
8
Experts active / token
2
Size this model
Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.
Size this model →