GPU Explorer

LLM inference planning โ€” compare GPU generations by memory, bandwidth, and cost efficiency

Preset:

Vendor:

X Axis:

Y Axis:

Loading chart...

๐Ÿ’ก Top-right = high VRAM and throughput. These GPUs handle larger models and longer contexts.

Bubble size represents tokens-per-dollar efficiency. Larger bubbles = better cost efficiency (more tokens generated per dollar spent).

Throughput Index is a planning metric derived from memory bandwidth, VRAM, and architecture generation. It enables relative GPU comparison โ€” not exact model throughput.

Inference performance depends on model architecture (GQA vs MHA), sequence length, batching, and inference backend (vLLM, TensorRT-LLM, etc.).

Hardware cost shown in data is GPU purchase price (USD, one-time). Hourly cloud pricing varies by provider and region.