Lenovo
Model: ThinkSystem SR680a V3
Request Your Server Quote
Tell us your workload and deployment needs. Our specialist replies within 24 hours via email, WhatsApp, or call.
Why this server is quoted to order
These servers are configured and quoted to order. Your build, storage, networking, warranty, and rack integration determine the final price, and your delivery destination sets shipping and customs. Submit the form below and our specialist will reply within 24 hours with a full quote including hardware, warranty, and worldwide DDP delivery.
Thanks! Our specialist will reply within 24 hours with your custom quote.
How your server order works
Submit form
Share workload & deployment details
Get your quote
Complete pricing within 24 hours
Review with specialist
Finalise configuration and delivery
Delivery
Shipped and ready for deployment
Genuine
Tested hardware
Worldwide
Global shipping
Support
Mining experts
Lenovo's 8U air-cooled platform for the NVIDIA HGX H200 baseboard: eight H200 SXM GPUs at the full 700W power class, joined by NVSwitch into 1.1TB of HBM3e per node, hosted by dual 5th Gen Intel Xeon Scalable processors with up to 2TB of TruDDR5 memory. Frontier-class training capability that deploys in a standard air-cooled data center, with no liquid cooling retrofit. Configured, quoted, and shipped worldwide DDP by MillionMiner.
Pricing, lead time, and hosting options. Personal advice from our sales team.
Pricing, lead time, and hosting options. Personal advice from our sales team.
This server is Lenovo's 8U air-cooled platform built around the NVIDIA HGX H200 8-GPU baseboard, a dual-Intel host engineered specifically so that the highest-power configuration of NVIDIA's Hopper generation can run in an ordinary air-cooled data center. In a market where full-power H200 nodes increasingly assume liquid cooling, that single engineering decision defines who this machine is for: organizations that want current-generation training capability without rebuilding their facility to get it. The GPU complex. Eight NVIDIA H200 SXM GPUs at the 700W power class mount to the HGX baseboard and interconnect through NVSwitch in a full mesh, every GPU reaching every other GPU at 900 GB/s without touching the CPU or PCIe path. Each GPU carries 141GB of HBM3e, the first HBM3e deployment in NVIDIA's lineup, giving the node 1,128GB of pooled GPU memory and aggregate bandwidth approaching 38 TB/s. FP8 throughput with the Transformer Engine exceeds 30 petaFLOPS per node. In workload terms: 70B-parameter models fine-tune at full FP16 precision within a single node, long-context inference runs with the KV cache headroom that 80GB-class GPUs cannot offer, and Mixture-of-Experts architectures route between experts at NVLink speed. Each H200 also partitions into as many as seven MIG instances at 16.5GB, so a single server can present more than fifty hardware-isolated GPU slices for multi-tenant serving. The thermal engineering, which is the reason to pick this chassis. Eight 700W GPUs plus dual server CPUs put a five-figure wattage of heat into one box. Lenovo's design answers it with 8U of airflow volume and fifteen hot-swap fans in an N+1 redundant arrangement, holding the full-power baseboard at specification on air alone. The practical consequence is large: no coolant distribution units, no facility plumbing, no liquid-loop maintenance contracts, and no constraint to the minority of data centers already fitted for direct liquid cooling. Power is matched to the same standard: eight 2600W hot-swap redundant supplies provide the envelope with failover headroom, drawing from ordinary data center AC distribution. The host platform. Dual 5th Gen Intel Xeon Scalable processors anchor the node, with configurations in the 48-core class per socket. Thirty-two DIMM slots carry TruDDR5 5600MHz memory to 2TB, which matters because the working rule for a training node is system memory at or above total GPU memory, and 2TB over a 1.1TB GPU pool clears it properly. Storage runs through sixteen hot-swap 2.5-inch PCIe 5.0 NVMe bays, enough local dataset capacity that most training jobs never reach across the network for data. Connectivity scales to 400 Gb/s per port, supporting the one-adapter-per-GPU topology that GPUDirect RDMA clustering requires when a single node becomes several. One chassis, two generations. The platform accepts either the HGX H200 baseboard or the HGX H100 baseboard with 80GB HBM3 GPUs. That flexibility is a genuine procurement lever: teams whose models fit comfortably in 640GB per node can specify H100 and bank the difference, while teams pushing memory limits take the 141GB H200. MillionMiner quotes both configurations on this platform. Where it sits in this catalog, stated plainly. The three 8x A100 systems here, the DGX A100, the Supermicro AS-4124GO-NART+, and the Exeton Quasar 640X, deliver proven NVSwitch training at the most accessible economics, 640GB per node. This Lenovo platform is the generational step up: 76 percent more GPU memory, 43 percent more bandwidth per GPU, FP8 Transformer Engine throughput the Ampere generation does not have, and NVLink at 900 GB/s against 600. The buyer it fits is the team whose workloads have outgrown A100 memory or whose training timelines justify Hopper throughput. Against the NVIDIA DGX H200, this is the OEM path: a configurable host, Lenovo's global service organization and XClarity management behind it, and three years of warranty coverage, without the appliance premium of NVIDIA's own box. Against the ASUS HGX H200 system also in this catalog, the Lenovo case is enterprise fleet standardization: organizations already running Lenovo server infrastructure keep one management plane, one support relationship, and one operational playbook. Every unit is quoted to configuration, tested, and shipped worldwide DDP with duties and customs handled. Deployment planning and hosting in MillionMiner's own data centers are available for teams that prefer not to provision roughly ten kilowatts of rack power on-site.
The decision that stalls most H200 deployments is not budget. It is facilities. The H200 SXM at its full 700W power class puts roughly 5,600W of GPU heat into one chassis, and much of the market answers that with direct liquid cooling, which means plumbing, coolant distribution, and a facility retrofit before the first training run. Lenovo's answer in this platform is different: an 8U chassis with fifteen hot-swap N+1 fans and the thermal volume to run the full-power HGX H200 baseboard on standard air. If your data center has power and conventional cooling, it can host this machine. What the node delivers once it is racked. Eight NVIDIA H200 SXM GPUs carry 141GB of HBM3e each, 1,128GB across the node, with NVSwitch joining every GPU to every other at 900 GB/s. That is the fabric that makes eight GPUs train like one large machine: gradient synchronization stays off the PCIe bus entirely, and scaling stays close to linear. Aggregate memory bandwidth approaches 38 TB/s, with FP8 Transformer Engine throughput beyond 30 petaFLOPS per node. The host side is sized to keep those GPUs fed. Dual 5th Gen Intel Xeon Scalable processors handle preprocessing and data loading. Up to 2TB of TruDDR5 5600MHz memory sits comfortably above the 1.1TB GPU pool, clearing the system-memory rule that undersized hosts violate. Sixteen hot-swap PCIe 5.0 NVMe bays keep datasets local at full speed, and network connectivity up to 400 Gb/s per port supports one fabric adapter per GPU for multi-node GPUDirect RDMA. The same chassis also accepts the HGX H100 baseboard, so procurement can choose the generation that fits the workload. Quoted and delivered worldwide DDP by MillionMiner.
Our mining specialists can help you find the perfect miner for your setup and budget.
This is Lenovo's platform for the NVIDIA HGX H200 8-GPU baseboard, and its defining trait is thermal: eight H200 SXM GPUs at the full 700W power class, NVSwitch-connected at 900 GB/s each, cooled entirely by air. The node pools 1.1TB of HBM3e GPU memory behind dual 5th Gen Intel Xeon Scalable processors, up to 2TB of TruDDR5 5600MHz memory, sixteen hot-swap PCIe 5.0 NVMe bays, and networking up to 400 Gb/s per port for GPUDirect RDMA clustering. Eight redundant 2600W power supplies and fifteen N+1 fans carry it. Frontier training without a liquid cooling retrofit, from a tier-one OEM.
Fifteen N+1 fans and 8U of thermal volume run the full-power HGX H200 baseboard without liquid cooling. No facility retrofit, no plumbing, no CDU.
Eight H200 SXM GPUs at 141GB each, every GPU linked at 900 GB/s. Fine-tune 70B models at full FP16 inside a single node.
One Lenovo chassis accepts either Hopper baseboard. XClarity management, global service coverage, and a configurable host behind the GPUs.
NVIDIA
$10,000.00
Supermicro
Contact for price
ASUS
Contact for price
Gigabyte
Contact for price
The specification set, an 8U air-cooled chassis with dual 5th Gen Intel Xeon Scalable processors, TruDDR5 memory, fifteen N+1 fans, and eight 2600W supplies carrying an HGX H100 or H200 baseboard, defines this Lenovo 8U air-cooled GPU platform. MillionMiner confirms the exact Lenovo model designation and configuration on your quote, since Lenovo offers this GPU complex across air-cooled Intel, air-cooled AMD, and liquid-cooled variants.
The chassis accepts either baseboard, and the decision is a memory decision. The H100 carries 80GB of HBM3 per GPU, 640GB per node, and remains excellent for models that fit within it. The H200 carries 141GB of HBM3e per GPU, 1.1TB per node, with 43 percent more bandwidth, and earns its premium when you fine-tune 70B-class models at full precision, serve long-context inference, or run memory-bound Mixture-of-Experts work. If your workloads already strain 80GB GPUs, the H200 is the answer; if not, the H100 configuration banks the difference.
Yes, and that is the platform's defining engineering. Fifteen hot-swap fans in an N+1 redundant arrangement move air through 8U of chassis volume, holding the full-power HGX baseboard at specification without direct liquid cooling. The practical meaning: any data center with adequate power and conventional cooling can host this machine, with no coolant distribution units, plumbing, or liquid-loop maintenance.
1,128GB of pooled HBM3e across eight GPUs, aggregate memory bandwidth approaching 38 TB/s, FP8 Transformer Engine throughput beyond 30 petaFLOPS, and 900 GB/s NVLink between every GPU pair through NVSwitch. In workload terms: full-precision fine-tuning of 70B-parameter models within one node, long-context inference with real KV cache headroom, and up to seven MIG partitions per GPU for multi-tenant serving.
Same class of GPU complex: eight H200 SXM GPUs on an NVSwitch baseboard. The DGX is NVIDIA's sealed appliance with its software stack and single-vendor support at appliance economics. This Lenovo platform is the OEM path: a configurable host with your choice of memory, storage, and networking, Lenovo's global service organization and XClarity management behind it, typically at a meaningfully lower platform cost. Certainty-led buyers tend toward the DGX; fleet-led and value-led enterprises tend here.
Both carry the same NVIDIA HGX H200 8-GPU baseboard, so GPU performance is equivalent. The decision is the host and the vendor relationship. The Lenovo case is enterprise standardization: organizations already operating Lenovo server fleets keep one management plane, one support contract, and one spares strategy. MillionMiner quotes both and will recommend based on your existing infrastructure rather than preference.
When memory or throughput has become the constraint. The A100 systems deliver 640GB per node and remain the value path for fine-tuning and serving within that envelope. The H200 node carries 1.1TB, 76 percent more memory and 43 percent more bandwidth per GPU, plus FP8 Transformer Engine throughput the Ampere generation lacks. Teams training larger models, serving longer contexts, or compressing training timelines are the ones for whom the step pays.
Yes. Connectivity scales to 400 Gb/s per port, supporting the one-adapter-per-GPU topology that GPUDirect RDMA requires, where GPUs in different nodes exchange gradients directly across the fabric without CPU involvement. Sixteen local NVMe bays keep datasets on-node, and MillionMiner advises on fabric and switch design when a deployment grows past one machine.
This is a roughly ten-kilowatt-class machine: eight 2600W hot-swap redundant supplies feed eight 700W GPUs plus the dual-CPU host, in an 8U rack footprint with high-volume front-to-back airflow. It requires data center power distribution but not liquid cooling, which is precisely its advantage. MillionMiner confirms the power and thermal plan for your site during configuration, and hosting in MillionMiner's own facilities is available as an alternative.
Submit your workload, GPU generation preference, and deployment details through the quote form. A MillionMiner specialist confirms the configuration, host memory, storage, networking, and the three-year warranty coverage, then the system is tested and shipped worldwide DDP with duties and customs handled before delivery. Rack integration guidance and hosted deployment are both available.