GPU Sizing Studio

Size LLM inference infrastructure — model weights, KV cache, multi-GPU topology, and vector database selection — across AWS, Azure, and GCP, then hand clients a branded report.

Already have an organization account? Sign in.

Weights + KV cache sizing
Dense and MoE architectures, FP32 → INT4 precision, GQA/MQA-aware KV cache math, paged-attention and quantized-KV options.
Multi-cloud GPU recommendations
Ranked AWS / Azure / GCP instance options with tensor-parallel sizing, multi-region pricing, and throughput/TTFT estimates.
Vector DB selection
Managed, self-hosted, and cloud-native options scored against your scale, QPS, latency, budget, and cloud ecosystem.

Free vs. Pro vs. Organization

Individuals sign up with just an email. Organizations are a separate, admin-provisioned workspace for teams — see an admin to get an account.

Free
$0
  • Core calculator, no account needed
  • Shareable result URL
  • Model directory — always public
Try the calculator
Free account
$0
  • Everything in Free
  • Save sizing runs
  • PDF/JSON export (1 per run)
  • Provenance footnotes
Create a free account
Pro
$29/mo
  • Everything in Free account
  • Quantization comparison view
  • Multi-node sizing
  • Full history + team-sharing links
Create a free account

Organizations get the full multi-client project workspace shown above (branded reports, capability matrices, live-fleet monitoring) via an admin-provisioned account — sign in if your team already has one.