DeepSeek

How much VRAM does DeepSeek-R1 need?

DeepSeek-R1 has 671B parameters, so its weights alone take about 1340 GB of VRAM at FP16 — or roughly 335 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.

Uses Multi-head Latent Attention (MLA), which compresses KV cache far below standard GQA math. This tool's KV estimate uses the generic GQA formula as a conservative upper bound. Reasoning models also produce long chain-of-thought output tokens — size avgOutputTokens generously.

Total params
671B
Active params
37B
Layers
61
Hidden size
7168
Attention heads
128
KV heads (GQA)
128
Vocab size
129,280
Native context window
131,072
Native precision
FP8
Experts (total)
256
Experts active / token
8
Size DeepSeek-R1 for your workload

Opens the free calculator with DeepSeek-R1 loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.

Size this model →