In Stock

NVIDIA

NVIDIA H100 SXM5 80GB HBM3 Hopper Tensor Core GPU for AI Training

Model: H100 SXM

Select Availability

Quantity

Total Price

$26,500.00

Buying 10 or more? Get custom bulk pricing

Pay with: Bank Transfer | BTC | ETH | USDT

Genuine

Tested hardware

Worldwide

Global shipping

Support

Mining experts

The NVIDIA H100 SXM 80GB is the standard every AI accelerator is measured against, built for distributed training at scale: the full GH100 Hopper die with 16,896 CUDA cores, 528 fourth-gen Tensor Cores and 80GB of HBM3 at 3.35 TB/s. NVLink 4.0 delivers 900 GB/s per GPU with NVSwitch connecting eight GPUs per HGX node, seven times PCIe bandwidth for near-linear scaling. 67 TFLOPS FP32 and 3,958 TFLOPS FP8 at up to 700W. MillionMiner configures, delivers and optionally hosts it worldwide.

Full Specifications

Model H100 SXM

Request a Bitcoin Miner Hosting Quote

Free quote, reply in 24h. No sales call.

4.4
star star star star star

4.7 / 5 on Trustpilot

Verified customer reviews

30,000+ miners delivered

Shipped worldwide since 2020

1,200+ customers globally

Trusted in 50+ countries

iso made-in-germany trustpilot
google-review

Get a Quote for the NVIDIA H100 SXM5 80GB HBM3 Hopper Tensor Core GPU for AI Training

Pricing, lead time, and hosting options. Personal advice from our sales team.

Reply within 24h via email, WhatsApp, or call.

Product Details

NVIDIA H100 SXM5 80GB Tensor Core GPU: Full Hopper Specifications, NVSwitch Architecture, and Enterprise AI Training Deployment

The H100 SXM exists because distributed AI training has a bandwidth problem that PCIe cannot solve. Training a 70B parameter model across multiple GPUs requires each GPU to exchange gradient updates with every other GPU after every forward and backward pass. On PCIe Gen 5 at 128 GB/s, those gradient exchanges become the bottleneck long before the GPUs run out of compute capacity. NVLink 4.0 at 900 GB/s per GPU (7x PCIe) and NVSwitch connecting all 8 GPUs in a node at full bandwidth eliminates that bottleneck. This is why every serious AI training deployment runs SXM, not PCIe. Full GH100 die specifications. 80 billion transistors at TSMC 4nm. 16,896 CUDA cores across 132 SMs (full die enablement on SXM versus 114 SMs on PCIe). 528 fourth-gen Tensor Cores supporting FP64, TF32, FP16, BF16, FP8, and INT8 with the Transformer Engine. 132 third-gen RT Cores. 80GB HBM3 on 5,120-bit bus at 3,350 GB/s bandwidth. 50MB L2 cache. The SXM5 form factor delivers approximately 30 percent more TFLOPS than the PCIe variant (67 versus 51 TFLOPS FP32) due to higher clock speeds enabled by the 700W power budget and HGX thermal infrastructure. Compute throughput at every precision tier. FP64: 34 TFLOPS (67 TFLOPS Tensor). FP32: 67 TFLOPS. TF32 Tensor: 989 TFLOPS with sparsity. FP16/BF16 Tensor: 1,979 TFLOPS with sparsity. FP8 Tensor: 3,958 TFLOPS with sparsity. INT8 Tensor: 3,958 TOPS with sparsity. The FP8 figure is the one that matters for transformer training: 3,958 TFLOPS with automatic precision management via the Transformer Engine means the H100 SXM delivers approximately 4x the training throughput of an A100 SXM on GPT-class models. NVLink and NVSwitch architecture. Each H100 SXM connects to the NVSwitch fabric via 18 NVLink 4.0 links providing 900 GB/s bidirectional bandwidth. NVSwitch provides all-to-all connectivity: any GPU can communicate with any other GPU in the same node at full 900 GB/s without routing through a CPU or PCIe bus. An 8-GPU HGX H100 node delivers 7.2 TB/s of aggregate NVLink bandwidth across all GPUs. For multi-node scaling, NVIDIA Quantum-2 NDR InfiniBand at 400 Gb/s per port extends the fabric beyond single nodes. DGX H100 versus HGX H100. DGX H100 is NVIDIA's turnkey 8-GPU system ($250,000 to $400,000 class) including CPUs, memory, storage, networking, and software stack. HGX H100 is the GPU baseboard module that server OEMs (Supermicro, Dell, HPE, Lenovo) integrate into their own server platforms. Both use the same 8x H100 SXM GPU configuration with NVSwitch. The HGX path offers more flexibility on CPU, storage, and networking choices. The SXM versus NVL versus PCIe decision framework. SXM (this product): maximum per-GPU performance, 8-GPU NVSwitch scaling, 700W TDP, requires HGX baseboard, optimized for distributed training. NVL (H100 NVL 94GB, separate MillionMiner listing): PCIe paired cards, 94GB unified memory, fits standard servers, optimized for large-model inference. PCIe (standard H100 80GB PCIe): single card at 350W, standard server slots, lower cost, limited to 2-GPU NVLink, suitable for single-GPU inference and fine-tuning. Choose SXM when training throughput and multi-GPU scaling efficiency are the priority. Choose NVL or PCIe when inference or infrastructure simplicity matters more. MIG on the H100 SXM creates up to 7 isolated instances at 10GB each. The most common production patterns per Spheron: 7x 1g.10gb for multi-tenant inference of small models, or 2x 3g.40gb for two simultaneous 13B model servers. Each MIG instance appears as a separate GPU device to the OS with hardware-enforced isolation. Confidential computing via Trusted Execution Environment (TEE) protects data and model weights during processing. This is a hardware-level security feature for compliance-sensitive AI deployments in healthcare (HIPAA), finance (SOC 2), and government (FedRAMP) where data cannot be exposed to the infrastructure operator. Up to 700W TDP configurable. Requires liquid cooling or high-airflow server chassis engineering. Standard air cooling is insufficient for sustained 700W operation. NVIDIA DGX H100 uses direct liquid cooling. HGX H100 configurations from Supermicro and Lenovo offer both air and liquid options depending on thermal budget.

NVIDIA H100 SXM5 80GB: The Multi-GPU Training Standard with 900 GB/s NVLink and 3,350 GB/s HBM3

The H100 SXM is the GPU that every other AI accelerator is benchmarked against. When NVIDIA, Google, Meta, Microsoft, and OpenAI publish training benchmarks, they run on H100 SXM clusters. When cloud providers quote AI compute capacity, they measure it in H100 SXM equivalents. This is the reference hardware for the current generation of AI. What separates the SXM from the PCIe variant is interconnect and power. NVLink 4.0 provides 900 GB/s bidirectional bandwidth per GPU, connecting up to 8 H100 SXM GPUs via NVSwitch in a single DGX or HGX node. That 900 GB/s is 7x faster than PCIe Gen 5 (128 GB/s) and enables near-linear scaling on distributed training workloads where gradient synchronization between GPUs is the bottleneck. The PCIe H100 tops out at 2-GPU NVLink pairs. The SXM scales to 8-GPU nodes and beyond through multi-node InfiniBand clusters. The full GH100 die runs at up to 700W TDP (configurable), delivering 16,896 CUDA cores and 528 fourth-gen Tensor Cores with the FP8 Transformer Engine. 80GB HBM3 at 3,350 GB/s bandwidth feeds those cores without memory starvation on large batch training. The Transformer Engine automatically manages FP8/FP16 mixed precision per neural network layer, delivering 4x training throughput over the A100 on transformer architectures without code changes. MIG creates up to 7 isolated instances at 10GB each for multi-tenant inference. Confidential computing (TEE) protects data and models during processing for compliance-sensitive deployments in healthcare, finance, and government. The SXM5 form factor requires an HGX baseboard (NVIDIA HGX H100 or DGX H100 platform). It does not plug into standard PCIe slots. This is purpose-built infrastructure for organizations committed to multi-GPU training at scale.

Need Help Choosing?

Our mining specialists can help you find the perfect miner for your setup and budget.

NVIDIA H100 SXM5 80GB HBM3 Tensor Core GPU

The GPU that defined the AI training era. Full GH100 Hopper die with 16,896 CUDA cores, 528 fourth-gen Tensor Cores, FP8 Transformer Engine, and 80GB HBM3 at 3,350 GB/s bandwidth. SXM5 mezzanine form factor for HGX baseboards. NVLink 4.0 at 900 GB/s per GPU with NVSwitch fabric connecting up to 8 GPUs per node. 67 TFLOPS FP32, 3,958 TFLOPS FP8 with sparsity. Up to 700W TDP. MIG for 7 isolated instances. Purpose-built for distributed AI training where inter-GPU bandwidth determines scaling efficiency.

900 GB/s NVLink 4.0 with NVSwitch

8 GPUs at full bandwidth in one node. 7.2 TB/s aggregate. Near-linear scaling on distributed training. The interconnect PCIe cannot match.

3,958 TFLOPS FP8 Transformer Engine

Fourth-gen Tensor Cores with automatic FP8/FP16 precision per layer. 4x training throughput over A100 on transformer architectures.

80GB HBM3 at 3,350 GB/s Bandwidth

68 percent faster memory bandwidth than A100. Feeds 16,896 CUDA cores without starvation on large batch training workloads.

FAQ

Frequently Asked Questions

SXM: full GH100 die at 16,896 CUDA cores, 700W TDP, 3,350 GB/s HBM3 bandwidth, NVLink 4.0 at 900 GB/s with NVSwitch connecting up to 8 GPUs. Requires HGX baseboard. PCIe: partially disabled die at 14,592 CUDA cores, 350W TDP, 2,000 GB/s bandwidth, NVLink limited to 2-GPU pairs via bridge. Fits standard servers. SXM is for distributed training at scale. PCIe is for inference and single-GPU workloads in existing infrastructure.

An NVIDIA HGX H100 baseboard or DGX H100 system. The SXM5 module does not plug into standard PCIe slots. It connects via the SXM5 mezzanine interface on the HGX baseboard. 700W TDP per GPU (5,600W for 8 GPUs) requires liquid cooling or enterprise-grade high-airflow chassis. HGX platforms are available from Supermicro, Dell, HPE, and Lenovo.

NVSwitch provides all-to-all GPU connectivity within a node. Each H100 SXM connects via 18 NVLink 4.0 links at 900 GB/s bidirectional. Any GPU communicates with any other GPU at full bandwidth without routing through CPU or PCIe. An 8-GPU node delivers 7.2 TB/s aggregate NVLink bandwidth. This is what enables near-linear scaling on distributed training where gradient synchronization between GPUs is the bottleneck.

Large-scale training of transformer models: GPT-class LLMs (70B to 175B+ parameters), vision transformers, multimodal models, and diffusion models. An 8-GPU HGX node with 640GB combined HBM3 handles training of 70B models with data parallelism and 175B+ with model parallelism. For inference, a single H100 SXM serves 70B models at FP8 quantization or 30B models at FP16.

H100 SXM: 67 TFLOPS FP32, 3,958 TFLOPS FP8, 80GB HBM3 at 3,350 GB/s, NVLink 4.0 at 900 GB/s, 700W. A100 SXM: 19.5 TFLOPS FP32, no FP8 support, 80GB HBM2e at 2,039 GB/s, NVLink 3.0 at 600 GB/s, 400W. The H100 delivers approximately 3x to 4x faster training on transformer models from the combined effect of higher bandwidth, FP8 precision, and faster NVLink.

Hardware-level automatic mixed-precision management. The Transformer Engine dynamically selects FP8 or FP16 precision per neural network layer during training and inference, maximizing throughput while maintaining model accuracy. This is a hardware feature unique to Hopper (H100) and newer architectures that requires no code changes from the developer.

Multi-Instance GPU creates up to 7 hardware-isolated instances at 10GB each. Each instance gets dedicated CUDA cores, Tensor Cores, L2 cache, and HBM with guaranteed QoS. Common production patterns: 7x 1g.10gb for multi-tenant inference of small models, or 2x 3g.40gb for two simultaneous 13B model servers. Each instance appears as a separate GPU device to the OS.

Hardware-based Trusted Execution Environment (TEE) that protects data and model weights during GPU processing. The infrastructure operator cannot access the data being computed. Required for compliance-sensitive AI deployments in healthcare (HIPAA), finance (SOC 2), and government (FedRAMP) where data privacy during processing is mandated.

Same Hopper architecture. The H200 SXM upgrades memory to 141GB HBM3e at 4,800 GB/s (versus 80GB HBM3 at 3,350 GB/s on the H100). Same CUDA core count and compute TFLOPS. The H200 is a memory and bandwidth upgrade for workloads constrained by HBM capacity or bandwidth on the H100, particularly larger model inference and training with bigger batch sizes.

NVIDIA H100 GPUs are subject to US export controls on advanced AI hardware. Not available in China, Hong Kong, and Macau. NVIDIA created the H800 (bandwidth-limited variant) for those markets. Confirm export eligibility with MillionMiner for your delivery destination before ordering.