Gemma
How much VRAM does Gemma 2 9B need?
Gemma 2 9B has 9.24B parameters, so its weights alone take about 18 GB of VRAM at FP16 — or roughly 5 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
Total params
9.24B
Active params
9.24B
Layers
42
Hidden size
3584
Attention heads
16
KV heads (GQA)
8
Vocab size
256,128
Native context window
8,192
Native precision
BF16
Size Gemma 2 9B for your workload
Opens the free calculator with Gemma 2 9B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →