Phi

Phi-3 Mini 3.8B

No GQA (MHA) — KV cache is proportionally larger per token than GQA models of similar size.

Total params
3.82B
Active params
3.82B
Layers
32
Hidden size
3072
Attention heads
32
KV heads (GQA)
32
Vocab size
32,064
Native context window
131,072
Native precision
BF16
Size this model

Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.

Size this model →