LLM GPU & VRAM calculator

Pick a model and see how much GPU memory it really needs — weights, KV cache, and a cloud instance that fits. It’s free and there’s no signup. Create a free account when you want to export a run or see where every number comes from.

Model

MoE: all 16 experts resident in memory; 1 active per token for compute. 10M-token context window is exceptional — most serving stacks size for a much smaller practical window.

Precision & KV cache

Allocates KV cache in fixed blocks instead of one contiguous slab, cutting wasted memory by roughly 30% — standard in vLLM.

Workload

Cloud target