Free, no-signup calculator — sign up only to save results

Right-size LLM infra before you buy it

Model weights, KV cache, multi-GPU topology, and vector database selection — sized across AWS, Azure, and GCP in minutes, then handed to clients as a branded report.

0
models in the catalog
0
GPU types, T4 to Blackwell
0
instance types across 6 clouds
0
vector databases scored

Everything one sizing pass needs

Three calculations, done together, so the numbers stay consistent with each other.

Weights + KV cache sizing
Dense and MoE architectures, FP32 → INT4 precision, GQA/MQA-aware KV cache math, paged-attention and quantized-KV options.
Multi-cloud GPU recommendations
Ranked AWS / Azure / GCP instance options with tensor-parallel sizing, pricing refreshed from each cloud, and throughput/TTFT estimates.
Vector DB selection
Managed, self-hosted, and cloud-native options scored against your scale, QPS, latency, budget, and cloud ecosystem.

Beyond the first calculation

Validate the call, keep it right after launch, and ship it as something a team can run.

Before you commit budget, see the trade-offs — precision, partitioning, commitment term, build vs. buy.

Quantization comparison
FP16 through INT4 side by side — VRAM saved, GPU tier unlocked, and the quality trade-off.
Self-host vs. API pricing
Self-hosted GPU cost per token lined up against managed APIs for chat, vision, OCR, transcription, and embeddings.
MIG-aware partitioning
NVIDIA Multi-Instance GPU profiles and a real placement solver, so compute and memory slices actually fit together.
Reserved & committed pricing
On-demand, 1-year, and 3-year committed-use rates modeled per cloud — not just list price.

From model name to GPU order

No account needed for the first three steps.

1
Pick a model
Search the directory or enter custom parameter counts and architecture details.
2
Set your workload
Context length, concurrency, precision, and latency targets for your use case.
3
Compare clouds
Ranked AWS, Azure, and GCP instances with pricing, throughput, and TTFT side by side.
4
Export the report
Save the run and hand a client or your team a branded, citable PDF.

Start free. Upgrade when it earns its keep.

The calculator is free for anyone. Pro adds exports, saved runs, and comparison tools. Teams sizing for clients get a shared workspace with projects, branded reports, and fleet auditing.

Free
$0
For a quick answer
  • Full calculator, no account
  • Shareable result link
  • Every model in the directory
Open the calculator
Most popular
Pro
See pricing
For serious individual use
  • PDF and JSON export
  • Saved run history
  • Quantization comparison
  • Team-sharing links
Create a free account
Teams
Talk to us
For consultants & platform teams
  • Multi-client projects & capability matrix
  • One-click deploy configs
  • Live fleet audits & waste alerts
  • Roles and audit log
See the team workspace

Teams workspaces are admin-provisioned, not self-serve — request access from your Pro account dashboard, or sign in if your team already has one.

Frequently asked questions

How accurate are the GPU and cost estimates?

Figures are engineering estimates derived from published architecture specs and vendor spec-sheet peak throughput and bandwidth. They're a strong starting point for sizing decisions, but we recommend validating against a benchmark before committing to a production purchase.

Do I need an account to use the calculator?

No. The core calculator and model directory are free and public — no account required. Creating a Pro account adds saved run history, PDF/JSON export, and quantization comparisons.

What's the difference between Pro and a Teams workspace?

Pro is a self-serve individual plan — sign up and upgrade with just an email. Teams is a separate, admin-provisioned workspace built for consultants and platform teams sizing for multiple clients: multi-client projects, branded reports, deploy configs, and live-fleet monitoring. Request access from your Pro account dashboard, or reach out and we'll set one up.

Can it tell me if something I've already deployed is over- or under-provisioned?

Yes — that's the right-sizing check. Import from a live fleet or enter what you're running and what it actually sees in production, and it flags waste against your configured risk thresholds, with alerts when something crosses a threshold.

Can I compare self-hosting against API providers?

Yes — self-hosted GPU cost per token is lined up against managed APIs for chat, vision, OCR, transcription, and embeddings, side by side, so build-vs-buy has real numbers behind it.

Do you generate the actual deployment files?

Yes — from a saved run or the capability matrix, generate ready-to-use systemd units, Docker Compose, Kubernetes manifests, or Terraform for AWS, Azure, or GCP, including vLLM flags and MIG device mappings where relevant.

Does this cover training as well as inference?

Both, as separate calculations. Inference sizing covers weights, KV cache, and ongoing serving cost; training/fine-tuning sizing answers a different question — the memory and one-time cost to finish a training run.

Which clouds and precisions are supported?

AWS, Azure, and GCP instance recommendations, with tensor-parallel sizing and pricing refreshed directly from each cloud (on-demand and 1/3-year committed). Precision ranges from FP32 down to INT4, including GQA/MQA-aware KV cache math, paged attention, and quantized-KV options.

Can I upgrade or cancel later?

Yes — Pro is billed monthly and can be changed or canceled at any time from account settings. There's no long-term commitment.

Size your next deployment in minutes

Free to try, no credit card, no account required to get a number.