Independent · 20 providers · verified 2026-07-06 · no vendor influence

The Live GPU Price Index

The cheapest place to rent every major NVIDIA GPU — H100 to B200 — ranked across 20 clouds, with on-demand and spot rates last verified 2026-07-06. The same card costs up to 10× more depending on where you rent it. We do the reading so you stop overpaying.

The Hyperscaler Premium

A hyperscaler charges up to 4.4× the specialist-cloud floor for the identical A100 80GB $3.43/hr at AWS versus $0.78/hr at Thunder Compute. Same silicon — you're paying for the platform around it.

Per-model premiums are in the "vs hyperscaler" column below.

Cheapest cloud GPU, by model

Lowest tracked on-demand rate per GPU, plus the memory-normalised cost ($/hr per 100GB VRAM) so you can compare cards on like terms. Tap a model for the full provider breakdown.

GPUVRAMCheapest $/hr$/hr · 100GBBest providerProvidersSpread
B200192GB$4.99$2.60Lambda52.9×
H200141GB$2.60$1.84GMI Cloud21.7×
H10080GB$1.65$2.06Vast.ai188.6×
A100 80GB80GB$0.78$0.97Thunder Compute67.4×
L40S48GB$0.72$1.50Spheron21.2×
RTX 409024GB$0.35$1.46Vast.ai32.0×
RTX 509032GB$0.86$2.69Spheron11.0×

37 price points across 20 providers · standard on-demand, per-GPU, USD · verified 2026-07-06.

Compare every provider

NVIDIA H100

80GB · 18 providers tracked · verified 2026-07-06

Cheapest on-demand

$1.65/hr

ProviderTypeOn-demand $/hrSpot $/hr
Vast.aimarketplacecheapest
PCIe
$1.65
$0.91Rent →
Spheron
SXM
$2.01
$1.43Rent →
UpCloud
SXM
$2.08
Rent →
Sesterce
SXM
$2.09
Rent →
FluidStack
SXM
$2.10
Rent →
CUDO Compute
SXM
$2.25
Rent →
Lambda
SXM
$2.49
Rent →
Novita AI
SXM
$2.59
Rent →
RunPod
SXM
$2.69
Rent →
Nebius
SXM
$2.95
Rent →
OVHcloud
PCIe
$2.99
Rent →
Vultr
SXM
$2.99
Rent →
Gcore
SXM
$3.21
Rent →
DigitalOcean
SXM
$3.39
Rent →
Paperspace
SXM
$5.95
Rent →
AWS
SXM
$6.88
Rent →
Microsoft Azure
SXM
$6.98
Rent →
Google Cloud
SXM
$14.19
Rent →

Standard published on-demand pricing, USD per single GPU per hour, last verified 2026-07-06. Spot/marketplace and committed-use rates run lower. Hyperscaler rates are per-GPU from multi-GPU instance list prices. Spread on H100: 8.6× between cheapest and dearest tracked rate.

Use this first

Pick the GPU by workload, then pick the provider by risk.

The cheapest row is not always the right row. Fine-tuning, long-context inference, batch jobs, and interactive notebooks fail in different ways. Start with the memory you need, then compare on-demand, spot, provider maturity, and how painful an interruption would be.

Memory-normalised view

Cheapest $/hr can hide expensive memory.

A 24GB card can look cheap until the model needs 80GB or 141GB. That is why the table includes $/hr per 100GB VRAM. In the current dataset, A100 80GB is the lowest memory-normalised on-demand option we track.

Breakeven warning

A rented GPU bills while idle.

A GPU kept at 30% utilisation costs more than three times its headline rate per useful hour. If your traffic is bursty, a managed API can still be cheaper even when the raw token math suggests self-hosting.

Why the same GPU costs 10× more elsewhere

An H100 is the same silicon whether you rent it from a specialist cloud or a hyperscaler — but the hourly price isn't. Specialist and marketplace providers compete on raw price; hyperscalers bundle the GPU with their platform, support, and networking and charge several times more. For a training run or a busy inference fleet, that gap is the difference between a healthy and a ruinous compute bill.

Renting vs. paying per token

Renting a GPU only beats a managed LLM API above a certain volume — and the line moves once you count the hours the card sits idle. Before you commit to a GPU, run your numbers through the self-host vs API breakeven.

How to read the index

Treat this page as the sourcing shortlist, not the final procurement decision. For the NVIDIA H100, the index compares 18 provider offers because it is still the default accelerator for many training and inference workloads. For the NVIDIA B200, the spread is smaller because public supply is newer and concentrated among fewer providers. That difference matters: a mature card gives you more provider choice, while a newer card may give you better throughput but less pricing transparency.

On-demand pricing is the cleanest way to compare providers because it is available without an auction, checkpointing strategy, or long-term commitment. Spot and marketplace rates are useful when your job can restart safely; they are dangerous for interactive sessions, customer-facing inference, or training runs without frequent checkpoints. Hyperscaler rates sit at the other end of the trade-off: higher $/hr, but stronger networking, compliance, support, and account controls.

Methodology

Rates are standard published on-demand pricing, per single GPU, in USD, last verified 2026-07-06. Spot and marketplace rates are shown where tracked and run lower with variable availability. Hyperscaler per-GPU figures are derived from multi-GPU instance list prices. Published under CC BY 4.0 — cite freely with a link. Spotted a stale rate? Tell us and we correct within 48 hours.