Quick estimate

Start with just a model name. We fill the rest, then let you tune every assumption.

✓ In catalog
Popular models: Llama 3.1, Mistral, Qwen 2.5, Gemma 2 — type to autocomplete

Stored in this browser only — never sent to our servers.

Based on your configuration — ISL 1000, OSL 150, FP16 KV cache, 97 concurrent users.
GPUs required0Configure workload below to see results↻ see math
How we got 0
total memory = 20 GB
usable / GPU = 72 GB (90% of 80)
⌈20 ÷ 72⌉ = 1 GPU
peak 3× → range up to 2
↻ flip back
Weight memory
0Nemotron-Mini-4B-Instruct↻ see math
Weight memory
params × bytes/param
8B × 2 (BF16)
= 16 GB
↻ flip back
KV cache / req
01150 tokens/req · 97 users↻ see math
KV cache / request
2 × layers × kv_heads ×
head_dim × bytes × tokens
32×8×128×2 = 128 KB/tok
× 150 tokens = 19 MB
↻ flip back
MONTHLY COST
CLOUD
$0.0K/mo
AWS · $0K over 5yr
SELF-HOSTED
$0.0K/mo
5yr amort · $0K total
Self-hosted saves $0.0K/mo
↻ see math
How we calculated this
Cloud:
0 GPUs × $1.22/gpu-hr × 730 hrs
= $0.0K/mo

Self-hosted:
$0K ÷ 60 months
= $0.0K/mo
(hardware amortization only)

5-year totals:
Cloud: $0K
Hardware: $0K

⚠️ Self-hosted excludes: power (~$X/mo), cooling, staff, networking. Typical full TCO adds 40–80% to this number.
↻ flip back
Want to change assumptions?

Every number above comes from these. Open a section to tune it — closed sections show their current values.

TP replicas

max_num_seqs chunked