GPU Explorer
LLM inference planning โ compare GPU generations by memory, bandwidth, and cost efficiency
Preset:
Vendor:
X Axis:
Y Axis:
Loading chart...
๐ก Top-right = high VRAM and throughput. These GPUs handle larger models and longer contexts.
Bubble size represents tokens-per-dollar efficiency. Larger bubbles = better cost efficiency (more tokens generated per dollar spent).
Throughput Index is a planning metric derived from memory bandwidth, VRAM, and architecture generation. It enables relative GPU comparison โ not exact model throughput.
Inference performance depends on model architecture (GQA vs MHA), sequence length, batching, and inference backend (vLLM, TensorRT-LLM, etc.).
Hardware cost shown in data is GPU purchase price (USD, one-time). Hourly cloud pricing varies by provider and region.