Buy NVIDIA GPUs at MillionMiner, from consumer Blackwell to data-center class, for AI and professional workloads. Available cards include the RTX 5090 32GB and RTX 4090 24GB in blower variants, the RTX PRO 6000 Blackwell 96GB, and the A100 and H100 Tensor Core GPUs. New with manufacturer warranty unless a listing states otherwise.
Every GPU ships free worldwide DDP, so the checkout price is the final landed cost. For multi-card builds, see our pre-built GPU servers, or have the cards deployed into our US hosting with power and cooling handled. Shop the full GPU range below.
Verified Stock
Every GPU tested & inspected before shipping
Fast Dispatch
Ships within 1–3 business days
Crypto Accepted
Pay with BTC, ETH, USDT & more
Expert Support
AI & GPU infrastructure specialists on hand
The AI revolution is powered by a remarkably small set of silicon chips. NVIDIA's data centre GPU lineup, from the H100 and H200 Tensor Core GPUs to the RTX PRO 6000 Blackwell and the RTX 5090, represents the compute substrate on which large language models, diffusion models, scientific simulations, and high-frequency inference pipelines run. We stock the full professional NVIDIA range — data centre accelerators, workstation GPUs, and professional visualisation cards — for enterprise buyers, research labs, AI startups, and cloud infrastructure operators.
H200 Memory BW
4.8 TB/s
HBM3e @ 141 GB capacity
H100 AI Performance
3.9 PFLOPS
FP8 sparse tensor operations
RTX PRO 6000 VRAM
96 GB
GDDR7 ECC — Blackwell architecture
RTX 5090 VRAM
32 GB GDDR7
Blackwell GB202 — 3352 GB/s BW
80 GB HBM2e
The workhorse of enterprise AI training. 3.35 TFLOPS FP64, NVLink 900 GB/s. SXM variant preferred for multi-GPU NVSwitch fabrics in DGX H100 systems.
141 GB HBM3e
Successor to H100. Same Hopper GPU die but 1.76× more memory and 4.8 TB/s bandwidth — critical for LLM inference with long context windows.
40 GB / 80 GB HBM2e
Previous-generation workhorse still in heavy use. 312 TFLOPS FP16 Tensor Core. Widely available in original and custom PCIe configurations.
96 GB GDDR7 ECC
The most capable workstation GPU ever built. Full PCIe form factor, NVLink bridge support, 4th-gen Tensor Cores. Ideal for local LLM inference and rendering.
32 GB GDDR7
NVIDIA's consumer flagship on Blackwell architecture. 21,760 CUDA cores, 3352 GB/s memory bandwidth. Best single-GPU for mixed workloads and inference.
24 GB GDDR6X
Previous Ada generation. 16,384 CUDA cores, 1008 GB/s BW. Still a strong performer and more accessible price point for inference clusters.
8 GB GDDR6
Low-profile professional GPU for edge inference and display-out workloads. Minimal power draw, silent operation, rack-dense deployment.
16 / 24 GB GDDR6
Inference-optimised Ada and Ampere PCIe GPUs. Designed for power-constrained server slots — 72 W TDP — with strong INT8 throughput for deployment at scale.
NVIDIA has shipped three major GPU microarchitectures since the AI boom began. The Hopper architecture (H100, H200) introduced the Transformer Engine — a hardware block that dynamically switches precision between FP8 and FP16 within a single layer pass to maximise throughput without meaningful accuracy loss. This was the design breakthrough that made training 70B+ parameter models commercially viable.
Ada Lovelace (RTX 4090, L4, L40S) brought the 4th-generation Tensor Core to the professional and prosumer market, alongside significantly improved ray tracing hardware. Ada cards occupy the sweet spot between consumer accessibility and professional capability, making them the dominant choice for boutique inference deployments and creative AI workstations.
Blackwell (RTX 5090, RTX PRO 6000, GB200 NVL72) is the current generation. Blackwell introduced a dual-die GPU design for the flagship data centre parts, NVLink 5.0 at 1.8 TB/s die-to-die bandwidth, and fifth-generation Tensor Cores with FP4 support. The RTX PRO 6000 Blackwell brings 96 GB ECC memory to the workstation tier — more capacity than a H100 PCIe — enabling previously cloud-exclusive workloads to run locally.
The three GPUs that define modern AI infrastructure at the workstation and single-node server tier.
The gold standard for distributed training. NVSwitch-connected SXM variant is mandatory for 8-GPU DGX-class nodes. PCIe variant fits standard servers.
Same Hopper GPU as H100 but nearly double the memory capacity and far greater bandwidth. Ideal for serving 70B+ parameter models at scale with long context lengths.
More VRAM than an H100 in a standard PCIe slot. Fits any workstation or tower server. Runs 70B models locally. The RTX PRO 6000 Server Edition removes display outputs for rack-optimised deployment.
The GPU that trains a model is not necessarily the best GPU to serve it. Training requires maximum FLOPS and large memory to hold parameters, gradients, and optimiser states simultaneously — a 70B model fine-tuned with AdamW requires over 500 GB of VRAM across a cluster. For training at this scale, H100 or H200 SXM nodes connected over NVSwitch are the industry standard, offering the combination of raw compute and the NVLink bandwidth needed to keep all GPUs fed.
Inference workloads care less about FLOPS and more about memory capacity and bandwidth. A quantised 70B model at INT4 fits in roughly 40 GB — meaning an H200 can serve it with room to spare, and a pair of RTX PRO 6000 Blackwell cards connected via NVLink Bridge can do the same. For high-throughput inference services, the H200's 4.8 TB/s bandwidth keeps latency low even under concurrent request batches.
Local workstation use cases — research, creative AI, local LLM chat, image generation, video synthesis — are well served by the RTX PRO 6000 Blackwell (96 GB), the RTX 5090 (32 GB), or the A100 (40/80 GB) depending on model size requirements. The RTX 4090 remains the most price-competitive option for users running quantised models under 24 GB.
Multi-node NVSwitch cluster. Mandatory for large-scale pre-training and full fine-tuning.
High-bandwidth memory essential for long context serving. NVLink Bridge for workstation pairs.
32–80 GB sufficient for quantised serving. Strong INT8 throughput.
Most VRAM in a PCIe card. Runs 34B models unquantised. ECC for reliability.
High CUDA core count + fast GDDR7 bandwidth ideal for diffusion inference.
Sub-75 W TDP. Single-slot or low-profile. Dense rack deployment.
ECC memory, certified drivers, large frame buffers for 3D/CAD/simulation.
The GB200 pairs two Blackwell GPU dies on a single package connected by NVLink-C2C at 1.8 TB/s — eliminating the PCIe bottleneck between processing units entirely.
New FP4 precision mode delivers up to 2× the throughput of FP8 for inference workloads. Maintains accuracy through NVIDIA's microscaling format (MXFP4).
Scale-up bandwidth doubles vs Hopper — 1.8 TB/s bidirectional per GPU. Critical for the NVL72 rack configuration tying 72 Blackwell GPUs into a single logical unit.
A dedicated Reliability, Availability & Serviceability engine for predictive hardware fault detection — reduces unplanned downtime in production inference clusters.
Hardware-isolated TEE (Trusted Execution Environment) for GPU workloads. Enables running proprietary models and sensitive inference in shared cloud infrastructure securely.
96 GB GDDR7 ECC in standard PCIe dual-slot form factor. Workstation-grade reliability with data-centre-class memory capacity. Server Edition removes display outputs.
The GeForce RTX 5090 launches Blackwell into the consumer and prosumer market. Built on the GB202 die with 21,760 CUDA cores and 32 GB of GDDR7 memory running at 1,792 GB/s bandwidth, it is the most powerful single-GPU that fits into a standard desktop PC. For AI researchers and developers who need local inference and experimentation without cloud costs, the RTX 5090 is the pragmatic choice.
The RTX PRO 6000 Blackwell takes a fundamentally different approach: rather than maximising consumer gaming performance, NVIDIA engineered a workstation GPU with 96 GB of ECC GDDR7 memory — more than any previous PCIe GPU from any manufacturer. The Server Edition and Max-Q Workstation Edition variants cover rack-mount and mobile deployment respectively. Two RTX PRO 6000 cards connected via NVLink Bridge give 192 GB of pooled memory — sufficient for running 70B parameter models unquantised.
For studios, research labs, and AI startups that cannot or will not run every workload in the cloud, the Blackwell workstation line changes the calculus entirely. Proprietary models stay on-premises. Inference latency drops. Data sovereignty is maintained. The upfront hardware cost is amortised against the cloud inference bills that would otherwise compound indefinitely.
Browse our full range of professional NVIDIA GPUs above, from single workstation cards to multi-GPU server configurations. Whether you're training your first model or scaling an inference cluster, our team will help you match the right hardware to your workload and budget.