Supermicro
Model: SYS-821GE-TNHR
Supermicro's 8U SuperServer built around the NVIDIA HGX H800 8-GPU baseboard: eight H800 SXM GPUs with 640GB of combined HBM3, full Hopper FP8 Transformer Engine throughput, and NVSwitch interconnect, hosted by dual… Intel Xeon Scalable processors with up to 8TB of DDR5 memory. The Hopper-generation compute that trained frontier models, on a Supermicro platform with air or optional liquid cooling. Quoted and shipped worldwide DDP by MillionMiner.
Request Your Server Quote
Tell us your workload and deployment needs. Our specialist replies within 24 hours via email, WhatsApp, or call.
Why this server is quoted to order
These servers are configured and quoted to order. Your build, storage, networking, warranty, and rack integration determine the final price, and your delivery destination sets shipping and customs. Submit the form below and our specialist will reply within 24 hours with a full quote including hardware, warranty, and worldwide DDP delivery.
Thanks! Our specialist will reply within 24 hours with your custom quote.
Prices and availability move with the market. We confirm both with you before your order becomes binding. Check price and availability
How your server order works
Submit form
Share workload & deployment details
Get your quote
Complete pricing within 24 hours
Review with specialist
Finalise configuration and delivery
Delivery
Shipped and ready for deployment
Genuine
Tested hardware
Worldwide
Global shipping
Support
Mining experts
Supermicro's 8U SuperServer built around the NVIDIA HGX H800 8-GPU baseboard: eight H800 SXM GPUs with 640GB of combined HBM3, full Hopper FP8 Transformer Engine throughput, and NVSwitch interconnect, hosted by dual… Intel Xeon Scalable processors with up to 8TB of DDR5 memory. The Hopper-generation compute that trained frontier models, on a Supermicro platform with air or optional liquid cooling. Quoted and shipped worldwide DDP by MillionMiner.
Pricing, lead time, and hosting options. Personal advice from our sales team.
Pricing, lead time, and hosting options. Personal advice from our sales team.
Most of what buyers read about the H800 is either dismissive or evasive. The accurate version is more useful: the H800 is the export-compliant variant of the H100, built on identical Hopper silicon, and the question of whether it fits your deployment has a clear answer once you know which two specifications were changed and whether your workload touches them. What is identical to the H100. The GPU die, the 80GB of HBM3 memory at full bandwidth, the fourth-generation Tensor Cores, and the complete precision stack that AI work actually uses: FP8 through the Transformer Engine, BF16, FP16, and TF32. An eight-GPU node delivers 640GB of pooled HBM3, aggregate memory bandwidth above 26 TB/s, and… FP8 throughput in the same class as an H100 node, beyond 30 petaFLOPS. MIG partitioning is intact as well, up to seven isolated instances per GPU, so one server can present more than fifty hardware-partitioned slices for multi-tenant inference. What was reduced, and when it matters. First, NVLink: 400 GB/s per GPU instead of 900. Inside a single node, NVSwitch still connects all eight GPUs in a full mesh, and 400 GB/s remains more than three times PCIe Gen 5 bandwidth, so single-node training, fine-tuning, and inference are barely touched. The reduction bites at multi-node frontier scale, where gradient synchronization across dozens of nodes leans on every gigabyte of fabric bandwidth. Second, FP64: cut drastically, which disqualifies the H800 for double-precision scientific simulation. AI training does not use FP64. If your work is computational fluid dynamics or climate modeling, buy an H100 system; if it is language models, vision, or recommendation, this reduction is irrelevant to you. The proof point, on the record. DeepSeek trained V3 and R1 on H800 clusters and published the engineering. Disciplined parallelism strategy on this exact GPU produced models that compete at the global frontier. That does not make the NVLink reduction disappear; it demonstrates that the H800's compute, memory, and FP8 throughput are genuinely frontier-class, and that the constraint is an engineering consideration rather than a ceiling. The Supermicro host. This is an 8U SuperServer with dual Intel Xeon Scalable processors on LGA-4677, up to 64 cores and 128 threads per socket. Memory scales to 8TB of ECC DDR5 across 32 DIMM slots, clearing the system-memory-above-GPU-memory rule many times over and leaving room for the in-memory preprocessing that large dataset pipelines want. Storage runs twelve hot-swap NVMe bays and three SATA bays for capacity, with two M.2 slots keeping the operating system off the dataset path. Networking ships with dual 10GbE onboard and optional 25GbE, with the PCIe expansion to add fabric adapters for GPUDirect RDMA when the deployment grows past one node. Power is six 3000W redundant supplies; cooling is ten heavy-duty fans on air, with direct liquid cooling available as an option for sites packing nodes densely. Supermicro's management stack, SuperCloud Composer and Supermicro Server Manager, handles fleet operation, with TPM 2.0 and Silicon Root of Trust underneath. Where it sits in this catalog. Against the three A100 systems, the DGX A100, the Supermicro AS-4124GO-NART+, and the Exeton Quasar 640X, this node is a full generation ahead where it counts: FP8 Transformer Engine throughput the Ampere generation does not have, HBM3 against HBM2e, and the same 640GB per node. Against the H100 and H200 systems, the Lenovo and ASUS HGX platforms, the H800 is the value position within the Hopper generation: single-node performance in the same class, with the NVLink and FP64 reductions priced in. The buyer it fits runs training, fine-tuning, or inference at one-to-few node scale and wants Hopper economics to reflect that. The buyer it does not fit runs FP64 simulation or multi-node training at cluster scale, and MillionMiner will say so in the quote conversation rather than after delivery. Export compliance note: the H800 is subject to US export controls, and MillionMiner confirms destination eligibility as part of every quote. Each system is configured to order, tested, and shipped worldwide DDP with duties and customs handled. Hosting in MillionMiner's own data centers is available for teams that prefer not to provision ten kilowatts of rack power on-site.
The H800 deserves a more honest introduction than it usually gets. It is not a cut-down afterthought. It is H100 silicon: the same Hopper architecture, the same 80GB of HBM3 at full memory bandwidth, the same fourth-generation Tensor Cores and FP8 Transformer Engine that define this GPU generation. Two things were reduced for export compliance: NVLink bandwidth, from 900 to 400 GB/s per GPU, and FP64 throughput, which matters for scientific simulation and almost nothing else in AI. Everything a training run or an inference fleet actually… exercises, FP8, BF16, TF32 compute and HBM3 bandwidth, is intact. The proof is public. DeepSeek trained V3 and R1, models that rearranged the global AI conversation, on clusters of exactly this GPU. The NVLink reduction is real, and it shows up in multi-node gradient synchronization at large cluster scale. Within a single node, where NVSwitch still connects all eight GPUs in a full mesh, fine-tuning, training, and inference run at Hopper pace. The Supermicro platform around the baseboard is dimensioned generously. Dual Intel Xeon Scalable processors with up to 64 cores each handle preprocessing and data loading. Thirty-two DIMM slots scale to 8TB of ECC DDR5, an order of magnitude above the 640GB GPU pool and far past the sizing rule that undersized hosts violate. Twelve NVMe bays plus three SATA bays keep datasets local, two M.2 slots isolate the boot path, and six 3000W redundant supplies carry the roughly ten-kilowatt envelope. Ten heavy-duty fans handle it on air, with liquid cooling available for dense rack deployments. Quoted to configuration and shipped worldwide DDP by MillionMiner.
Our mining specialists can help you find the perfect miner for your setup and budget.
The H800 is the GPU that proved a point: DeepSeek trained its V3 and R1 frontier models on H800 clusters, running the same Hopper silicon, 80GB of HBM3, and FP8 Transformer Engine as the H100. This Supermicro 8U SuperServer carries eight of them on an NVSwitch baseboard, 640GB of GPU memory per node, behind dual Intel Xeon Scalable processors, up to 8TB of DDR5 across 32 DIMM slots, twelve NVMe bays, and six redundant 3000W power supplies. Ten heavy-duty fans cool it on air, with liquid cooling optional. Hopper-class training and inference, quoted and shipped worldwide DDP by MillionMiner.
H800 clusters trained DeepSeek V3 and R1. Same Hopper silicon, 80GB HBM3, and FP8 Transformer Engine as the H100, proven at the frontier.
Eight H800 SXM GPUs on NVSwitch deliver 30+ petaFLOPS of FP8 and 26 TB/s of aggregate memory bandwidth. MIG splits each GPU seven ways.
Up to 8TB DDR5 across 32 DIMMs, twelve NVMe bays, six 3000W redundant supplies, ten fans on air with liquid cooling optional.
The H800 is the export-compliant variant of the H100, built on identical Hopper silicon with the same 80GB of HBM3 at full bandwidth and the same FP8 Transformer Engine. Two specifications were reduced: NVLink bandwidth, 400 GB/s per GPU instead of 900, and FP64 throughput, which is used by scientific simulation rather than AI. For training, fine-tuning, and inference, the compute and memory you actually exercise match the H100.
The public record answers this. DeepSeek trained V3 and R1, frontier models by any measure, on H800 clusters and published the engineering behind it. The GPU's FP8 throughput, HBM3 bandwidth, and Tensor Core architecture are H100-class because they are the same silicon. The NVLink reduction is an engineering consideration at multi-node scale, not a performance ceiling.
At multi-node cluster scale. Inside one server, NVSwitch still joins all eight GPUs in a full mesh at 400 GB/s each, more than three times PCIe Gen 5, so single-node training and fine-tuning are barely affected. Synchronizing gradients across dozens of nodes is where the gap against 900 GB/s H100 fabric widens. Buyers running one to a few nodes rarely feel it; buyers planning large clusters should weigh the H100 and H200 systems in this catalog, and MillionMiner will model both in the quote.
Not if the work is double-precision. The H800's FP64 throughput was cut drastically as part of export compliance, which disqualifies it for computational fluid dynamics, climate modeling, and similar FP64 workloads. For those, the H100 and H200 systems in this catalog keep full FP64 Tensor performance. For AI training and inference, which run FP8, BF16, and TF32, the reduction is irrelevant.
Eight H800 SXM GPUs with 640GB of pooled HBM3, aggregate memory bandwidth above 26 TB/s, FP8 Transformer Engine throughput beyond 30 petaFLOPS, and full-mesh NVSwitch interconnect. Each GPU partitions into up to seven MIG instances, so the node can present more than fifty hardware-isolated GPU slices for multi-tenant inference serving.
Dual Intel Xeon Scalable processors on LGA-4677, up to 64 cores and 128 threads per socket, with 32 DIMM slots scaling to 8TB of ECC DDR5. Storage runs twelve hot-swap NVMe bays plus three SATA bays, with two M.2 slots keeping the boot path separate from dataset I/O. Networking ships with dual 10GbE onboard, optional 25GbE, and PCIe expansion for fabric adapters. Supermicro's SuperCloud Composer and Server Manager handle fleet management, with TPM 2.0 and Silicon Root of Trust for platform security.
No. Ten heavy-duty fans cool the node on standard data center air, which keeps deployment inside ordinary facilities. Direct liquid cooling is available as an option for sites packing nodes at high rack density, and MillionMiner advises on which configuration fits your facility during the quote.
It is a full generation ahead where AI workloads notice. The Hopper FP8 Transformer Engine roughly doubles effective training throughput against Ampere at the same node memory, HBM3 outpaces the A100's HBM2e, and NVSwitch bandwidth per GPU is comparable to the A100's 600 GB/s fabric. Both deliver 640GB per node. Teams choosing between them are weighing Ampere economics against Hopper throughput, and MillionMiner quotes both honestly.
Single-node performance sits in the same class as an H100 system, since the silicon, memory, and FP8 throughput match. The H100 and H200 platforms justify their position with full 900 GB/s NVLink for multi-node scale, full FP64 for HPC, and in the H200's case 141GB of HBM3e per GPU. The H800 is the value position within the Hopper generation for one-to-few node AI deployments. Outgrow it, and the upgrade path is already in this catalog.
Submit your workload and deployment details through the quote form, and a MillionMiner specialist confirms configuration, destination eligibility under the US export controls that apply to the H800, and the delivery plan. Every system is tested before shipment and delivered worldwide DDP with duties and customs handled. Hosting in MillionMiner's own data centers is available as an alternative to on-site deployment.