Phi
How much VRAM does Phi-3 Mini 3.8B need?
Phi-3 Mini 3.8B has 3.82B parameters, so its weights alone take about 8 GB of VRAM at FP16 — or roughly 2 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
No GQA (MHA) — KV cache is proportionally larger per token than GQA models of similar size.
Total params
3.82B
Active params
3.82B
Layers
32
Hidden size
3072
Attention heads
32
KV heads (GQA)
32
Vocab size
32,064
Native context window
131,072
Native precision
BF16
Size Phi-3 Mini 3.8B for your workload
Opens the free calculator with Phi-3 Mini 3.8B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →