GPU Sizing Studio
Size LLM inference infrastructure — model weights, KV cache, multi-GPU topology, and vector database selection — across AWS, Azure, and GCP, then hand clients a branded report.
Already have an organization account? Sign in.
Weights + KV cache sizing
Dense and MoE architectures, FP32 → INT4 precision, GQA/MQA-aware KV cache math, paged-attention and quantized-KV options.
Multi-cloud GPU recommendations
Ranked AWS / Azure / GCP instance options with tensor-parallel sizing, multi-region pricing, and throughput/TTFT estimates.
Vector DB selection
Managed, self-hosted, and cloud-native options scored against your scale, QPS, latency, budget, and cloud ecosystem.
Free vs. Pro vs. Organization
Individuals sign up with just an email. Organizations are a separate, admin-provisioned workspace for teams — see an admin to get an account.
Free
$0
- Core calculator, no account needed
- Shareable result URL
- Model directory — always public
Free account
$0
- Everything in Free
- Save sizing runs
- PDF/JSON export (1 per run)
- Provenance footnotes
Pro
$29/mo
- Everything in Free account
- Quantization comparison view
- Multi-node sizing
- Full history + team-sharing links
Organizations get the full multi-client project workspace shown above (branded reports, capability matrices, live-fleet monitoring) via an admin-provisioned account — sign in if your team already has one.