Million Miner Logo
Hosting & Colocation · 18 min read · Aug 04, 2026 · Updated Aug 11, 2026

Cloud GPU Pricing in 2026: Rent vs Buy, and the Break-Even Math Nobody Shows You

Sara Chen

Bitcoin & Altcoin Analyst

Cloud GPU Pricing in 2026: Rent vs Buy, and the Break-Even Math Nobody Shows You
A survey of 25 cloud providers found the same NVIDIA H100 renting for anywhere between $0.80 and $11.10 an hour. Same silicon, same 80 GB, a 13.8x difference in price. That single fact explains why “what does a cloud GPU cost” has no one-line answer, and why people burn real money on both sides of it: renting hyperscaler capacity for workloads a marketplace would serve at a fifth the price, or buying $30,000 cards that sit at 20% utilization while a rental would have cost half as much.

We run a GPU cloud, sell the same hardware outright, and host customer-owned machines, which means we get paid under all three models and have no incentive to pretend one of them always wins. This guide is the honest map: the full 2026 price table by provider tier, the six factors behind the 13.8x spread, the break-even formula between renting and owning with worked numbers, the third option the rent-vs-buy binary ignores, and the cost traps that live in invoices rather than pricing pages.

The short answer

  • Typical mid-2026 on-demand rates: H100 from $1.49 to 2.69/hr on marketplaces and specialist clouds, $3.90 to 12.29/hr at hyperscalers; A100 from $0.68/hr; B200 from $3.75 to 16.11/hr; RTX 4090-class from $0.14/hr spot.
  • The rule that decides rent vs buy: sustained utilization. Below roughly 40%, renting usually wins. Above it, owning pulls ahead, and 24/7 inference makes ownership overwhelming.
  • Worked break-even: a ~$30,000 H100 against a $2.50/hr rental crosses over near 14,000 GPU-hours, roughly 19 months at full utilization, before resale value shortens it further.
  • The third path: buy the card, colocate it in a hosted facility. Ownership economics without running power, cooling, and uptime yourself.


The 2026 cloud GPU price table

Every number below is an on-demand range compiled from published provider price sheets and tracked indexes in July 2026. The columns matter more than the cells: the market is three different businesses wearing one name.
Cloud GPU price table for mid-2026: H100 from $1.49 to $12.29 per hour, A100 from $0.68, H200 from $2.30, B200 from $3.75 to $16.11, and RTX 4090 from $0.14 spot, across marketplace, specialist, hyperscaler, and spot tiers.
Marketplaces (Vast.ai, RunPod's community tier) aggregate thousands of third-party hosts and post the lowest stickers, H100s from $1.49, RTX 4090s from $0.14 spot, with reliability that floats host by host. Specialist clouds (Lambda, CoreWeave, RunPod secure) run their own fleets: H100s at roughly $1.99 to 3.99, the sweet spot where most serious workloads land. Hyperscalers (AWS, Azure, GCP) price the same card at $3.90 to $12.29 with enterprise SLAs, compliance, and whole-VM bundling, and even they are falling: AWS cut H100 rates about 44% in mid-2025. Two dynamics frame 2026: previous-generation cards keep sliding toward commodity pricing (A100s from $0.68, with analysts expecting sub-$1 broadly), while the newest silicon runs the other way, B200 and B300 listings roughly doubled in a year as constrained supply met hyperscaler premiums. Which card your workload actually needs is a separate question, answered by VRAM first and throughput second, which is what our VRAM guide and benchmark tool exist for.

Why the same H100 costs $0.80 or $11.10 an hour

Why H100 rental prices vary 13.8x: SLA and uptime guarantees, interconnect class (PCIe vs SXM with NVLink), commitment terms, whole-VM bundling, egress and storage fees, and host quality by region.
None of the spread is a scam; all of it is product difference hiding behind one GPU name. A $0.80 marketplace listing is a bare PCIe card on somebody's hardware with no uptime promise. An $11 hyperscaler instance is an SXM card with NVLink inside an 8-GPU machine, bundled CPU, RAM, and disk, a 99.9% SLA, and a compliance department. Between them sit the differences that actually price the market: interconnect class (a PCIe H100 and an NVLinked SXM H100 are different machines for multi-GPU work), commitment terms (reserved 1-to-12-month deals run 20 to 40% below on-demand, spot runs 40 to 65% below with interruption risk), and what is bundled versus billed separately. The practical rule: never compare prices across tiers, only within one, holding all six factors constant, and treat any comparison of a whole-VM price against a bare-card price as broken by construction.

Rent vs buy: the break-even math

Renting has no capex and no commitment; owning has a purchase price and a resale value. The crossover between them is one division away, and almost nobody publishes it with the assumptions attached.
GPU rent-vs-buy break-even: net purchase cost divided by rental rate minus owning cost per hour; a $30,000 H100 against $2.50/hr rental crosses near 14,000 GPU-hours, an RTX 4090 48GB near 4,700 hours, with resale value cutting effective break-even 30 to 40 percent.
Work the flagship case. An H100 costs roughly $25,000 to 40,000 new; take $30,000, add professional hosting at ~$0.35 per GPU-hour including power, and rent the same class at ~$2.50. Break-even lands near 14,000 GPU-hours: about 19 months at 100% utilization, four years at 40%. Now the two corrections that swing it. Resale is real money: data-center GPUs hold value (used A100s still trade at $4,000 to 9,000 years after launch), and a 40 to 50% residual after two years cuts effective break-even by roughly a third. Utilization is everything: at 20% duty cycle the purchase never pays back before obsolescence; at steady 24/7 inference it pays back inside two years and then prints. That is the whole decision, and it is why the 40% rule from our A100 vs H100 analysis generalizes: below ~40% sustained utilization rent, above it own. Your electricity rate moves the line too, since power is the dominant operating cost, the same economics our AI energy study maps at grid scale.

One honest asymmetry: renting lets you change your mind about the chip. Owning an H100 through the Rubin transition means riding its resale curve down; renting means switching next month. Price that flexibility, it is not free but it is not worthless.

Ready to Start Mining?

Free worldwide DDP shipping. Professional hosting from $0.055/kWh.

The third path: buy and colocate

The three GPU capacity paths compared: renting (elastic, perpetual cost), buying and self-hosting (cheapest but you run power, cooling, and uptime), and buying plus colocation (ownership economics with data-center operations handled).
The rent-vs-buy framing quietly assumes that buying means becoming your own data center, and that assumption does most of the work in rent-favoring conclusions, because power provisioning, cooling, networking, and 3 a.m. failures are genuinely expensive to self-operate. The path operators actually use splits the difference: buy the hardware, colocate it in a facility that sells contracted power, cooling, and uptime, the model behind our GPU hosting and the reason mining companies transitioned into AI infrastructure so naturally, since the facility skill set is identical. Colocation adds a hosting fee to the ownership column, the ~$0.35/hr in the worked example above, and in exchange removes every operational objection to owning. For steady workloads above the 40% line, it is usually the cheapest total-cost path available, which is exactly why we sell it, a bias you should weigh alongside the math.
A GPU cluster racked in a data center: whether rented by the hour or owned and colocated, this is the hardware every cloud GPU price ultimately meters.


Six cost traps the hourly rate hides

Six cloud GPU cost traps: whole-VM versus bare-card pricing, egress fees, idle billing on attached resources, spot interruption losses, minimum commitments behind headline rates, and availability as the real price.
The hourly rate is the advertised price; the invoice is the real one. The traps, in rough order of damage: comparing a hyperscaler's whole-VM price (CPU, RAM, and disk included) against a marketplace's bare card; egress fees on moving weights and datasets out, which several specialist clouds charge nothing for and hyperscalers charge plenty; idle billing on attached volumes and provisioned-but-stopped instances; spot interruptions wiping out uncheckpointed training runs, converting a 60% discount into a net loss; headline rates that turn out to be 1-to-12-month commitments with on-demand at twice the number; and the quiet one, availability: the cheapest tier being sold out in your region at your scale means the real price is the next tier up. Budget on obtainable capacity, verify what the training workload actually needs before reserving it, and read the billing page as carefully as the pricing page.

What this guide cannot decide for you

Three limits, stated plainly. First, these prices are July 2026 snapshots of a market that moves monthly, mid-generation cards drift down, newest cards spike on scarcity, and providers cut rates in double-digit percentages without warning, so treat every figure as a range to verify against live pages, including ours, before committing money. Second, the break-even math is directional planning, not a quote: your electricity rate, negotiated hosting fee, financing cost, and resale timing all move the crossover, and the worked examples state their assumptions so you can swap in your own. Third, we are a seller in this market three times over, cloud, hardware, and hosting, which is why every recommendation here reduces to a formula you can check rather than a claim you have to trust. Run your utilization honestly; it will make the decision for you.

The bottom line

Cloud GPU pricing in 2026 is a three-tier market with a 13.8x spread for identical silicon: marketplaces from $1.49/hr for an H100, specialists around $2 to 4, hyperscalers to $12, with previous-generation cards sliding toward commodity rates while B200-class scarcity holds premiums. The rent-vs-buy answer is a utilization question wearing a pricing costume: under ~40% sustained use, rent, ideally on the cheapest tier whose reliability your workload tolerates; over it, own, and let colocation remove the operational excuse; and always price the third path before accepting the binary. Match the card to the workload with the choosing guide, price ownership in the GPU shop, and whichever side of the break-even line you land on, make the line itself, not a pricing page headline, the thing that decides.

Frequently asked questions

How much does it cost to rent a GPU in the cloud?

Mid-2026 on-demand ranges: H100 80GB from $1.49 to 2.69/hr on marketplaces and specialist clouds and $3.90 to 12.29/hr at hyperscalers; A100 80GB from $0.68/hr; H200 from about $2.30/hr; B200 from $3.75 to 16.11/hr; RTX 4090-class consumer cards from $0.14/hr spot. Reserved commitments cut 20 to 40% off on-demand and spot instances 40 to 65%, with interruption risk. Prices move monthly, so verify live pages before budgeting.

How much does it cost to rent an H100 per hour?

From about $1.49/hr on GPU marketplaces (Vast.ai class), $1.99 to 3.99/hr on specialist clouds (RunPod, Lambda, CoreWeave class), and $3.90 to 12.29/hr at hyperscalers (AWS, Azure, GCP), as of mid-2026. A 25-provider survey found the full spread running $0.80 to $11.10 for identical silicon. The differences buy SLA, interconnect class, bundling, and compliance, not a faster chip.

Is it cheaper to rent or buy a GPU for AI?

It depends almost entirely on sustained utilization. Below roughly 40% duty cycle, renting is usually cheaper because idle owned hardware still cost its purchase price. Above 40%, owning pulls ahead: a ~$30,000 H100 against a $2.50/hr rental breaks even near 14,000 GPU-hours (about 19 months at full utilization), and resale value cuts that by roughly a third. Steady 24/7 inference makes ownership decisively cheaper.

Why are cloud GPU prices so different between providers?

Because the same GPU name covers different products. Marketplaces list bare cards on third-party hosts with no SLA; specialist clouds run owned fleets with support; hyperscalers bundle whole-VM resources, enterprise SLAs, and compliance. Add interconnect class (PCIe vs NVLinked SXM), commitment discounts of 20 to 40%, spot discounts of 40 to 65%, egress policies, and regional supply, and a 13.8x spread for identical silicon is the predictable result.

What is the cheapest way to run AI workloads on GPUs?

For spiky or experimental workloads: spot or community-marketplace instances with checkpointing, from $0.14/hr for RTX 4090-class cards and around $2.25/hr for H100 spot. For steady production workloads: owning hardware and colocating it in a hosted facility usually wins above roughly 40% utilization, combining ownership economics with professional power and uptime. The most expensive common mistake is running steady 24/7 inference on on-demand hyperscaler rates.

Are GPU rental prices going up or down?

Both, by generation. Previous-generation cards keep falling: AWS cut H100 pricing about 44% in mid-2025, specialist H100 rates now start under $2, and A100s trade from $0.68/hr with analysts expecting sub-$1 broadly. The newest silicon runs opposite: B200 and B300 on-demand listings roughly doubled over the past year as constrained supply met premium hyperscaler pricing. Expect each generation to follow the same arc: scarce and pricey, then abundant and commodity.

What hidden costs come with cloud GPUs?

The recurring ones: data egress fees on moving models and datasets out (zero on several specialist clouds, substantial on hyperscalers), storage and idle billing on attached volumes and stopped instances, whole-VM pricing that bundles CPU and RAM you did not price, and minimum commitments behind headline rates. The catastrophic one: spot interruptions wiping uncheckpointed training runs, which converts a 60% discount into a net loss. Read the billing documentation, not just the pricing page.

Does buying a GPU hold its value?

Better than most hardware. Data-center GPUs retain meaningful resale value years into life: used A100s still trade at $4,000 to 9,000, and H100s hold strong residuals while demand outruns supply. A realistic 40 to 50% residual after two years shortens the effective rent-vs-buy break-even by roughly a third. The risk is generational: each new architecture (B200 now, Rubin next) steepens the depreciation curve of the one before it.

What is GPU colocation and when does it make sense?

Colocation means you own the GPUs and a hosting facility runs them: contracted power, cooling, networking, and uptime for a monthly or hourly fee. It makes sense when your utilization sits above roughly 40% for the long term but you do not want to operate power and cooling yourself: you keep ownership economics and resale value while outsourcing the data-center problem. It is the standard path for steady inference workloads and the model our own hosting business runs on.

Should I rent a B200 or buy an H100 in 2026?

For most teams: neither extreme. B200 rentals carry scarcity premiums ($3.75 to 16.11/hr) and reservation queues, while buying an H100 near the Rubin transition means riding its steepening resale curve. The pragmatic 2026 plays: rent H100/H200 capacity on specialist clouds (under $3/hr) for flexible work, or buy and colocate H100-class or 48 to 96 GB workstation cards for steady inference above the 40% utilization line. Re-run the break-even when Rubin pricing lands.

Sources and notes
Price ranges compiled July 2026 from published provider price sheets (Lambda, RunPod, Vast.ai, CoreWeave, AWS, Azure, GCP) and tracked indexes including AIMultiple's 67-provider monthly GPU index; the 13.8x spread figure is from a 2026 community survey of 25 providers; the AWS ~44% H100 cut is from June 2025 pricing announcements. Purchase prices and the 40% utilization rule follow our A100 vs H100 analysis; used-market ranges reflect mid-2026 listings. Break-even examples state their assumptions in text and are directional planning figures, not quotes; hosting cost assumes professional colocation including power at typical contracted rates. We operate a GPU cloud, sell GPUs, and host customer hardware; this triple position is disclosed in the article. All prices move monthly; verify live pages before purchase decisions. Hero: U.S. Department of Energy, public domain; body: CSIRO, CC BY 3.0, via Wikimedia Commons. Diagrams: original MillionMiner graphics, free to reuse with attribution and a link to this page. Informational content, not financial advice

Ready to Start Mining?

Free worldwide DDP shipping. Professional hosting from $0.055/kWh.

Sara Chen

Written by

Sara Chen

Bitcoin & Altcoin Analyst

Sara covers network fundamentals, mining economics, and emerging proof-of-work protocols. She holds a background in applied mathematics and has tracked Bitcoin difficulty adjustments since block 600,000.

Comments 0

Please sign in to leave a comment

Delete Comment?

This action cannot be undone.