Buy GPUs for Sale: NVIDIA RTX 5090, A100, H100 and RTX PRO 6000

Buy NVIDIA GPUs at MillionMiner, from consumer Blackwell to data-center class, for AI and professional workloads. Available cards include the RTX 5090 32GB and RTX 4090 24GB in blower variants, the RTX PRO 6000 Blackwell 96GB, and the A100 and H100 Tensor Core GPUs. New with manufacturer warranty unless a listing states otherwise.

Every GPU ships free worldwide DDP, so the checkout price is the final landed cost. For multi-card builds, see our pre-built GPU servers, or have the cards deployed into our US hosting with power and cooling handled. Shop the full GPU range below.

5.0
star star star star star
4.97
star star star star star
4.5
star star star star star
honeycomb-grid honeycomb-grid
Filter & Sort

Verified Stock

Every GPU tested & inspected before shipping

Fast Dispatch

Ships within 1–3 business days

Crypto Accepted

Pay with BTC, ETH, USDT & more

Expert Support

AI & GPU infrastructure specialists on hand

AI & GPU Computing

Professional AI GPUs for Training, Inference & HPC Workloads

The AI revolution is powered by a remarkably small set of silicon chips. NVIDIA's data centre GPU lineup, from the H100 and H200 Tensor Core GPUs to the RTX PRO 6000 Blackwell and the RTX 5090, represents the compute substrate on which large language models, diffusion models, scientific simulations, and high-frequency inference pipelines run. We stock the full professional NVIDIA range — data centre accelerators, workstation GPUs, and professional visualisation cards — for enterprise buyers, research labs, AI startups, and cloud infrastructure operators.

H200 Memory BW

4.8 TB/s

HBM3e @ 141 GB capacity

H100 AI Performance

3.9 PFLOPS

FP8 sparse tensor operations

RTX PRO 6000 VRAM

96 GB

GDDR7 ECC — Blackwell architecture

RTX 5090 VRAM

32 GB GDDR7

Blackwell GB202 — 3352 GB/s BW


NVIDIA Professional GPU Lineup

From Workstation to Hyperscale Data Centre

H100 SXM / PCIe

80 GB HBM2e

The workhorse of enterprise AI training. 3.35 TFLOPS FP64, NVLink 900 GB/s. SXM variant preferred for multi-GPU NVSwitch fabrics in DGX H100 systems.

H200 SXM / PCIe

141 GB HBM3e

Successor to H100. Same Hopper GPU die but 1.76× more memory and 4.8 TB/s bandwidth — critical for LLM inference with long context windows.

A100 PCIe / SXM

40 GB / 80 GB HBM2e

Previous-generation workhorse still in heavy use. 312 TFLOPS FP16 Tensor Core. Widely available in original and custom PCIe configurations.

RTX PRO 6000 Blackwell

96 GB GDDR7 ECC

The most capable workstation GPU ever built. Full PCIe form factor, NVLink bridge support, 4th-gen Tensor Cores. Ideal for local LLM inference and rendering.

RTX 5090 (GB202)

32 GB GDDR7

NVIDIA's consumer flagship on Blackwell architecture. 21,760 CUDA cores, 3352 GB/s memory bandwidth. Best single-GPU for mixed workloads and inference.

RTX 4090 (Ada)

24 GB GDDR6X

Previous Ada generation. 16,384 CUDA cores, 1008 GB/s BW. Still a strong performer and more accessible price point for inference clusters.

T1000 (Professional)

8 GB GDDR6

Low-profile professional GPU for edge inference and display-out workloads. Minimal power draw, silent operation, rack-dense deployment.

A2 / L4 (PCIe)

16 / 24 GB GDDR6

Inference-optimised Ada and Ampere PCIe GPUs. Designed for power-constrained server slots — 72 W TDP — with strong INT8 throughput for deployment at scale.

The NVIDIA Architecture Story

Hopper, Ada Lovelace & Blackwell — Three Generations Driving the AI Era

NVIDIA has shipped three major GPU microarchitectures since the AI boom began. The Hopper architecture (H100, H200) introduced the Transformer Engine — a hardware block that dynamically switches precision between FP8 and FP16 within a single layer pass to maximise throughput without meaningful accuracy loss. This was the design breakthrough that made training 70B+ parameter models commercially viable.

Ada Lovelace (RTX 4090, L4, L40S) brought the 4th-generation Tensor Core to the professional and prosumer market, alongside significantly improved ray tracing hardware. Ada cards occupy the sweet spot between consumer accessibility and professional capability, making them the dominant choice for boutique inference deployments and creative AI workstations.

Blackwell (RTX 5090, RTX PRO 6000, GB200 NVL72) is the current generation. Blackwell introduced a dual-die GPU design for the flagship data centre parts, NVLink 5.0 at 1.8 TB/s die-to-die bandwidth, and fifth-generation Tensor Cores with FP4 support. The RTX PRO 6000 Blackwell brings 96 GB ECC memory to the workstation tier — more capacity than a H100 PCIe — enabling previously cloud-exclusive workloads to run locally.


Side-by-Side

H100 vs H200 vs RTX PRO 6000 Blackwell

The three GPUs that define modern AI infrastructure at the workstation and single-node server tier.

Hopper

H100 SXM5

Best for Training
Memory 80 GB HBM2e
Memory BW 3.35 TB/s
FP8 Sparse 3.9 PFLOPS
TDP 700 W (SXM)
NVLink NVLink 4 — 900 GB/s
Form Factor SXM5 / PCIe

The gold standard for distributed training. NVSwitch-connected SXM variant is mandatory for 8-GPU DGX-class nodes. PCIe variant fits standard servers.

Hopper

H200 SXM5

Best for Inference
Memory 141 GB HBM3e
Memory BW 4.8 TB/s
FP8 Sparse 3.9 PFLOPS
TDP 700 W (SXM)
NVLink NVLink 4 — 900 GB/s
Form Factor SXM5 / PCIe

Same Hopper GPU as H100 but nearly double the memory capacity and far greater bandwidth. Ideal for serving 70B+ parameter models at scale with long context lengths.

Blackwell

RTX PRO 6000 BW

Best Workstation
Memory 96 GB GDDR7 ECC
Memory BW ~1.8 TB/s
Tensor Cores 5th-gen, FP4/FP8/FP16
TDP 300 W (PCIe)
NVLink NVLink Bridge (optional)
Form Factor Dual-slot PCIe x16

More VRAM than an H100 in a standard PCIe slot. Fits any workstation or tower server. Runs 70B models locally. The RTX PRO 6000 Server Edition removes display outputs for rack-optimised deployment.


Choosing the Right GPU

Training vs Inference vs Workstation: Which GPU Do You Need?

The GPU that trains a model is not necessarily the best GPU to serve it. Training requires maximum FLOPS and large memory to hold parameters, gradients, and optimiser states simultaneously — a 70B model fine-tuned with AdamW requires over 500 GB of VRAM across a cluster. For training at this scale, H100 or H200 SXM nodes connected over NVSwitch are the industry standard, offering the combination of raw compute and the NVLink bandwidth needed to keep all GPUs fed.

Inference workloads care less about FLOPS and more about memory capacity and bandwidth. A quantised 70B model at INT4 fits in roughly 40 GB — meaning an H200 can serve it with room to spare, and a pair of RTX PRO 6000 Blackwell cards connected via NVLink Bridge can do the same. For high-throughput inference services, the H200's 4.8 TB/s bandwidth keeps latency low even under concurrent request batches.

Local workstation use cases — research, creative AI, local LLM chat, image generation, video synthesis — are well served by the RTX PRO 6000 Blackwell (96 GB), the RTX 5090 (32 GB), or the A100 (40/80 GB) depending on model size requirements. The RTX 4090 remains the most price-competitive option for users running quantised models under 24 GB.

Quick Selection Guide

Match Your Workload to the Right Hardware

LLM Training (70B+)
H100 / H200 SXM

Multi-node NVSwitch cluster. Mandatory for large-scale pre-training and full fine-tuning.

LLM Inference (70B)
H200 or 2× RTX PRO 6000

High-bandwidth memory essential for long context serving. NVLink Bridge for workstation pairs.

LLM Inference (7–13B)
RTX 5090 / A100 80GB

32–80 GB sufficient for quantised serving. Strong INT8 throughput.

Local AI Workstation
RTX PRO 6000 BW (96 GB)

Most VRAM in a PCIe card. Runs 34B models unquantised. ECC for reliability.

Image / Video Gen
RTX 5090 / RTX 4090

High CUDA core count + fast GDDR7 bandwidth ideal for diffusion inference.

Edge / Low-Power Inference
T1000 / L4 / A2

Sub-75 W TDP. Single-slot or low-profile. Dense rack deployment.

Professional Visualisation
RTX PRO 6000 / A100

ECC memory, certified drivers, large frame buffers for 3D/CAD/simulation.


Blackwell Architecture Highlights

What Makes Blackwell a Generational Leap

Dual-Die GPU Design

The GB200 pairs two Blackwell GPU dies on a single package connected by NVLink-C2C at 1.8 TB/s — eliminating the PCIe bottleneck between processing units entirely.

5th-Gen Tensor Cores

New FP4 precision mode delivers up to 2× the throughput of FP8 for inference workloads. Maintains accuracy through NVIDIA's microscaling format (MXFP4).

NVLink 5.0

Scale-up bandwidth doubles vs Hopper — 1.8 TB/s bidirectional per GPU. Critical for the NVL72 rack configuration tying 72 Blackwell GPUs into a single logical unit.

RAS Engine

A dedicated Reliability, Availability & Serviceability engine for predictive hardware fault detection — reduces unplanned downtime in production inference clusters.

Confidential Computing

Hardware-isolated TEE (Trusted Execution Environment) for GPU workloads. Enables running proprietary models and sensitive inference in shared cloud infrastructure securely.

RTX PRO 6000 — PCIe Flagship

96 GB GDDR7 ECC in standard PCIe dual-slot form factor. Workstation-grade reliability with data-centre-class memory capacity. Server Edition removes display outputs.

NVIDIA Blackwell

The RTX 5090 and RTX PRO 6000: Blackwell for Professionals

The GeForce RTX 5090 launches Blackwell into the consumer and prosumer market. Built on the GB202 die with 21,760 CUDA cores and 32 GB of GDDR7 memory running at 1,792 GB/s bandwidth, it is the most powerful single-GPU that fits into a standard desktop PC. For AI researchers and developers who need local inference and experimentation without cloud costs, the RTX 5090 is the pragmatic choice.

The RTX PRO 6000 Blackwell takes a fundamentally different approach: rather than maximising consumer gaming performance, NVIDIA engineered a workstation GPU with 96 GB of ECC GDDR7 memory — more than any previous PCIe GPU from any manufacturer. The Server Edition and Max-Q Workstation Edition variants cover rack-mount and mobile deployment respectively. Two RTX PRO 6000 cards connected via NVLink Bridge give 192 GB of pooled memory — sufficient for running 70B parameter models unquantised.

For studios, research labs, and AI startups that cannot or will not run every workload in the cloud, the Blackwell workstation line changes the calculus entirely. Proprietary models stay on-premises. Inference latency drops. Data sovereignty is maintained. The upfront hardware cost is amortised against the cloud inference bills that would otherwise compound indefinitely.


FAQ

Frequently Asked Questions

Both use the same GH100 Hopper GPU die and deliver identical compute FLOPS (~3.9 PFLOPS FP8 sparse). The H200 replaces HBM2e with HBM3e memory: capacity increases from 80 GB to 141 GB and bandwidth grows from 3.35 TB/s to 4.8 TB/s. The H200 is therefore superior for inference-heavy workloads where large models must fit on a single GPU and high throughput is required.

For inference workloads yes — its 96 GB ECC GDDR7 capacity exceeds an H100 PCIe's 80 GB and it fits in any standard PCIe workstation slot. For training, the H100/H200 SXM form factor with NVSwitch delivers significantly higher inter-GPU bandwidth and greater raw FLOPS at the same TDP. The RTX PRO 6000 is the better choice when you need high VRAM capacity locally without NVSwitch infrastructure.

SXM GPUs (H100 SXM5, H200 SXM5) mount to a specialised baseboard rather than a PCIe slot and connect to each other via NVSwitch fabric, enabling 900 GB/s of inter-GPU bandwidth per GPU. PCIe variants fit standard servers and workstations but are limited to NVLink Bridge (2 GPUs) or PCIe bandwidth (x16, ~64 GB/s) for multi-GPU communication. For clusters of 8+ GPUs requiring all-to-all communication, SXM NVSwitch systems are mandatory.

At FP16 (16-bit) precision a 70B model requires approximately 140 GB of VRAM — a single H200 (141 GB) handles this at the margin. At INT8 quantisation, approximately 70 GB is needed. At INT4 (GGUF Q4), approximately 35–40 GB, fitting within an H100 80 GB or a pair of RTX 4090s. The RTX PRO 6000 Blackwell (96 GB) handles INT8 70B models on a single GPU with room for KV cache.

The RTX 5090 is excellent for local AI inference, image generation (Stable Diffusion, Flux, DALL-E 3 replicas), video synthesis, and fine-tuning of small-to-medium models. Its 32 GB GDDR7 is sufficient for running 13B models in FP16 and 70B models at INT4. For multi-GPU training clusters, data centre GPUs remain preferable due to NVLink bandwidth and larger memory capacity.

The NVIDIA A100 (Ampere architecture, 2020) remains a capable and widely used AI GPU. It ships in 40 GB and 80 GB HBM2e configurations, delivering 312 TFLOPS FP16 Tensor Core performance. As supply has loosened relative to H100, A100 pricing is more accessible. For inference and moderate training workloads it remains strong. For cutting-edge research or the largest models, H100/H200 are now preferred.

Yes. MillionMiner serves enterprise, research, and cloud infrastructure customers with volume GPU orders. Contact our sales team for pricing on multi-unit H100, H200, A100, and RTX PRO 6000 purchases. We can arrange freight logistics, staging, and invoice in multiple currencies including crypto.

Our team is available 24/7 via WhatsApp (+49 176 777 888 33), email and phone. We help with GPU selection for training vs inference workloads, multi-GPU cluster design, and B2B procurement. Contact us or visit our FAQ for instant answers.

Ready to Build Your AI Infrastructure?

Browse our full range of professional NVIDIA GPUs above, from single workstation cards to multi-GPU server configurations. Whether you're training your first model or scaling an inference cluster, our team will help you match the right hardware to your workload and budget.