Independent · 20 providers · verified 2026-07-06 · no vendor influence
The Live GPU Price Index
The cheapest place to rent every major NVIDIA GPU — H100 to B200 — ranked across 20 clouds, with on-demand and spot rates last verified 2026-07-06. The same card costs up to 10× more depending on where you rent it. We do the reading so you stop overpaying.
The Hyperscaler Premium
A hyperscaler charges up to 4.4× the specialist-cloud floor for the identical A100 80GB — $3.43/hr at AWS versus $0.78/hr at Thunder Compute. Same silicon — you're paying for the platform around it.
Per-model premiums are in the "vs hyperscaler" column below.
Cheapest cloud GPU, by model
Lowest tracked on-demand rate per GPU, plus the memory-normalised cost ($/hr per 100GB VRAM) so you can compare cards on like terms. Tap a model for the full provider breakdown.
| GPU | VRAM | Cheapest $/hr | $/hr · 100GB | Best provider | Providers | Spread |
|---|---|---|---|---|---|---|
| B200 | 192GB | $4.99 | $2.60 | Lambda ↗ | 5 | 2.9× |
| H200 | 141GB | $2.60 | $1.84 | GMI Cloud | 2 | 1.7× |
| H100 | 80GB | $1.65 | $2.06 | Vast.ai ↗ | 18 | 8.6× |
| A100 80GB | 80GB | $0.78 | $0.97 | Thunder Compute | 6 | 7.4× |
| L40S | 48GB | $0.72 | $1.50 | Spheron | 2 | 1.2× |
| RTX 4090 | 24GB | $0.35 | $1.46 | Vast.ai ↗ | 3 | 2.0× |
| RTX 5090 | 32GB | $0.86 | $2.69 | Spheron | 1 | 1.0× |
37 price points across 20 providers · standard on-demand, per-GPU, USD · verified 2026-07-06.
Compare every provider
NVIDIA H100
80GB · 18 providers tracked · verified 2026-07-06
Cheapest on-demand
$1.65/hr
| Provider | Type | On-demand $/hr | Spot $/hr | |
|---|---|---|---|---|
Vast.aimarketplacecheapest | PCIe | $1.65 | $0.91 | Rent → |
Spheron | SXM | $2.01 | $1.43 | Rent → |
UpCloud | SXM | $2.08 | — | Rent → |
Sesterce | SXM | $2.09 | — | Rent → |
FluidStack | SXM | $2.10 | — | Rent → |
CUDO Compute | SXM | $2.25 | — | Rent → |
Lambda | SXM | $2.49 | — | Rent → |
Novita AI | SXM | $2.59 | — | Rent → |
RunPod | SXM | $2.69 | — | Rent → |
Nebius | SXM | $2.95 | — | Rent → |
OVHcloud | PCIe | $2.99 | — | Rent → |
Vultr | SXM | $2.99 | — | Rent → |
Gcore | SXM | $3.21 | — | Rent → |
DigitalOcean | SXM | $3.39 | — | Rent → |
Paperspace | SXM | $5.95 | — | Rent → |
AWS | SXM | $6.88 | — | Rent → |
Microsoft Azure | SXM | $6.98 | — | Rent → |
Google Cloud | SXM | $14.19 | — | Rent → |
Standard published on-demand pricing, USD per single GPU per hour, last verified 2026-07-06. Spot/marketplace and committed-use rates run lower. Hyperscaler rates are per-GPU from multi-GPU instance list prices. Spread on H100: 8.6× between cheapest and dearest tracked rate.
Use this first
Pick the GPU by workload, then pick the provider by risk.
The cheapest row is not always the right row. Fine-tuning, long-context inference, batch jobs, and interactive notebooks fail in different ways. Start with the memory you need, then compare on-demand, spot, provider maturity, and how painful an interruption would be.
Memory-normalised view
Cheapest $/hr can hide expensive memory.
A 24GB card can look cheap until the model needs 80GB or 141GB. That is why the table includes $/hr per 100GB VRAM. In the current dataset, A100 80GB is the lowest memory-normalised on-demand option we track.
Breakeven warning
A rented GPU bills while idle.
A GPU kept at 30% utilisation costs more than three times its headline rate per useful hour. If your traffic is bursty, a managed API can still be cheaper even when the raw token math suggests self-hosting.
Why the same GPU costs 10× more elsewhere
An H100 is the same silicon whether you rent it from a specialist cloud or a hyperscaler — but the hourly price isn't. Specialist and marketplace providers compete on raw price; hyperscalers bundle the GPU with their platform, support, and networking and charge several times more. For a training run or a busy inference fleet, that gap is the difference between a healthy and a ruinous compute bill.
Renting vs. paying per token
Renting a GPU only beats a managed LLM API above a certain volume — and the line moves once you count the hours the card sits idle. Before you commit to a GPU, run your numbers through the self-host vs API breakeven.
How to read the index
Treat this page as the sourcing shortlist, not the final procurement decision. For the NVIDIA H100, the index compares 18 provider offers because it is still the default accelerator for many training and inference workloads. For the NVIDIA B200, the spread is smaller because public supply is newer and concentrated among fewer providers. That difference matters: a mature card gives you more provider choice, while a newer card may give you better throughput but less pricing transparency.
On-demand pricing is the cleanest way to compare providers because it is available without an auction, checkpointing strategy, or long-term commitment. Spot and marketplace rates are useful when your job can restart safely; they are dangerous for interactive sessions, customer-facing inference, or training runs without frequent checkpoints. Hyperscaler rates sit at the other end of the trade-off: higher $/hr, but stronger networking, compliance, support, and account controls.
Methodology
Rates are standard published on-demand pricing, per single GPU, in USD, last verified 2026-07-06. Spot and marketplace rates are shown where tracked and run lower with variable availability. Hyperscaler per-GPU figures are derived from multi-GPU instance list prices. Published under CC BY 4.0 — cite freely with a link. Spotted a stale rate? Tell us and we correct within 48 hours.