Mistral
How much VRAM does Mistral Large (2411) need?
Mistral Large (2411) has 123B parameters, so its weights alone take about 245 GB of VRAM at FP16 — or roughly 62 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
Total params
123B
Active params
123B
Layers
88
Hidden size
12288
Attention heads
96
KV heads (GQA)
8
Vocab size
32,768
Native context window
131,072
Native precision
BF16
Size Mistral Large (2411) for your workload
Opens the free calculator with Mistral Large (2411) loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →