GPU sizing calculator
Free, no account needed. Get a free account for exports and accuracy footnotes, or go Pro for quantization comparison, history, and team sharing.
Model
MoE: all 16 experts resident in memory; 1 active per token for compute. 10M-token context window is exceptional — most serving stacks size for a much smaller practical window.
Precision & KV cache
Allocates KV cache in fixed blocks instead of one contiguous slab, cutting wasted memory by roughly 30% — standard in vLLM.