Right-size LLM infra before you buy it
Model weights, KV cache, multi-GPU topology, and vector database selection — sized across AWS, Azure, and GCP in minutes, then handed to clients as a branded report.
Everything one sizing pass needs
Three calculations, done together, so the numbers stay consistent with each other.
Beyond the first calculation
Validate the call, keep it right after launch, and ship it as something a team can run.
Before you commit budget, see the trade-offs — precision, partitioning, commitment term, build vs. buy.
From model name to GPU order
No account needed for the first three steps.
Start free. Upgrade when it earns its keep.
The calculator is free for anyone. Pro adds exports, saved runs, and comparison tools. Teams sizing for clients get a shared workspace with projects, branded reports, and fleet auditing.
- Full calculator, no account
- Shareable result link
- Every model in the directory
- PDF and JSON export
- Saved run history
- Quantization comparison
- Team-sharing links
- Multi-client projects & capability matrix
- One-click deploy configs
- Live fleet audits & waste alerts
- Roles and audit log
Teams workspaces are admin-provisioned, not self-serve — request access from your Pro account dashboard, or sign in if your team already has one.
Frequently asked questions
How accurate are the GPU and cost estimates?
Figures are engineering estimates derived from published architecture specs and vendor spec-sheet peak throughput and bandwidth. They're a strong starting point for sizing decisions, but we recommend validating against a benchmark before committing to a production purchase.
Do I need an account to use the calculator?
No. The core calculator and model directory are free and public — no account required. Creating a Pro account adds saved run history, PDF/JSON export, and quantization comparisons.
What's the difference between Pro and a Teams workspace?
Pro is a self-serve individual plan — sign up and upgrade with just an email. Teams is a separate, admin-provisioned workspace built for consultants and platform teams sizing for multiple clients: multi-client projects, branded reports, deploy configs, and live-fleet monitoring. Request access from your Pro account dashboard, or reach out and we'll set one up.
Can it tell me if something I've already deployed is over- or under-provisioned?
Yes — that's the right-sizing check. Import from a live fleet or enter what you're running and what it actually sees in production, and it flags waste against your configured risk thresholds, with alerts when something crosses a threshold.
Can I compare self-hosting against API providers?
Yes — self-hosted GPU cost per token is lined up against managed APIs for chat, vision, OCR, transcription, and embeddings, side by side, so build-vs-buy has real numbers behind it.
Do you generate the actual deployment files?
Yes — from a saved run or the capability matrix, generate ready-to-use systemd units, Docker Compose, Kubernetes manifests, or Terraform for AWS, Azure, or GCP, including vLLM flags and MIG device mappings where relevant.
Does this cover training as well as inference?
Both, as separate calculations. Inference sizing covers weights, KV cache, and ongoing serving cost; training/fine-tuning sizing answers a different question — the memory and one-time cost to finish a training run.
Which clouds and precisions are supported?
AWS, Azure, and GCP instance recommendations, with tensor-parallel sizing and pricing refreshed directly from each cloud (on-demand and 1/3-year committed). Precision ranges from FP32 down to INT4, including GQA/MQA-aware KV cache math, paged attention, and quantized-KV options.
Can I upgrade or cancel later?
Yes — Pro is billed monthly and can be changed or canceled at any time from account settings. There's no long-term commitment.
Size your next deployment in minutes
Free to try, no credit card, no account required to get a number.