GPU sizing calculator

Free, no account needed. Get a free account for exports and accuracy footnotes, or go Pro for quantization comparison, history, and team sharing.

Model

MoE: all 16 experts resident in memory; 1 active per token for compute. 10M-token context window is exceptional — most serving stacks size for a much smaller practical window.

Precision & KV cache

Allocates KV cache in fixed blocks instead of one contiguous slab, cutting wasted memory by roughly 30% — standard in vLLM.

Workload

Cloud target