Gemma
How much VRAM does Gemma 3 27B need?
Gemma 3 27B has 27B parameters, so its weights alone take about 54 GB of VRAM at FP16 — or roughly 14 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
Total params
27B
Active params
27B
Layers
62
Hidden size
5376
Attention heads
32
KV heads (GQA)
16
Vocab size
262,208
Native context window
131,072
Native precision
BF16
Size Gemma 3 27B for your workload
Opens the free calculator with Gemma 3 27B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →