DeepSeek

DeepSeek-V3

Uses Multi-head Latent Attention (MLA), which compresses KV cache far below standard GQA math. This tool's KV estimate uses the generic GQA formula as a conservative upper bound.

Total params
671B
Active params
37B
Layers
61
Hidden size
7168
Attention heads
128
KV heads (GQA)
128
Vocab size
129,280
Native context window
131,072
Native precision
FP8
Experts (total)
256
Experts active / token
8
Size this model

Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.

Size this model →