Gemma

How much VRAM does Gemma 2 27B need?

Gemma 2 27B has 27.2B parameters, so its weights alone take about 54 GB of VRAM at FP16 — or roughly 14 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.

Total params
27.2B
Active params
27.2B
Layers
46
Hidden size
4608
Attention heads
32
KV heads (GQA)
16
Vocab size
256,128
Native context window
8,192
Native precision
BF16
Size Gemma 2 27B for your workload

Opens the free calculator with Gemma 2 27B loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.

Size this model →