Supermicro

Supermicro AS-4124GO-NART+ 8x A100 HGX AI Server

Model: AS-4124GO-NART+

Request Your Server Quote

Tell us your workload and deployment needs. Our specialist replies within 24 hours via email, WhatsApp, or call.

Why this server is quoted to order

These servers are configured and quoted to order. Your build, storage, networking, warranty, and rack integration determine the final price, and your delivery destination sets shipping and customs. Submit the form below and our specialist will reply within 24 hours with a full quote including hardware, warranty, and worldwide DDP delivery.

How your server order works

1

Submit form

Share workload & deployment details

2

Get your quote

Complete pricing within 24 hours

3

Review with specialist

Finalise configuration and delivery

4

Delivery

Shipped and ready for deployment

Genuine

Tested hardware

Worldwide

Global shipping

Support

Mining experts

A true 8-GPU HGX A100 system: eight NVIDIA A100 80GB SXM4 GPUs on an NVSwitch baseboard delivering 640GB of pooled HBM2e and 600 GB/s GPU-to-GPU NVLink, hosted by dual AMD EPYC processors in a 4U chassis. DGX A100-class distributed training capability on a build-to-order platform, configured to your workload by MillionMiner.

Full Specifications

Model AS-4124GO-NART+
Form Factor 4U Rackmount
GPU 8x NVIDIA A100 80GB SXM4 (HGX baseboard)
GPU Interconnect NVSwitch full-mesh, 600 GB/s NVLink per GPU
GPU Memory 640GB pooled HBM2e (8x 80GB)
Processor Dual AMD EPYC 7002/7003 Series (PCIe Gen 4 platform)
Cores / Threads Up to 128 cores / 256 threads (dual socket)
Memory Up to 8TB DDR4-3200 ECC (32 DIMM slots) — configured to order
Storage Hot-swap NVMe bays — configured to order
Networking Up to 1 NIC per GPU for GPUDirect RDMA (InfiniBand / high-speed Ethernet) — configured to order
Power Supply Redundant 3000W Titanium Level

Request a Bitcoin Miner Hosting Quote

Free quote, reply in 24h. No sales call.

4.4
star star star star star

4.7 / 5 on Trustpilot

Verified customer reviews

30,000+ miners delivered

Shipped worldwide since 2020

1,200+ customers globally

Trusted in 50+ countries

iso made-in-germany trustpilot
google-review

Get a Quote for the Supermicro AS-4124GO-NART+ 8x A100 HGX AI Server

Pricing, lead time, and hosting options. Personal advice from our sales team.

Reply within 24h via email, WhatsApp, or call.

Product Details

Supermicro AS-4124GO-NART+ 8x A100 80GB HGX Server: NVSwitch Architecture, Configuration Guide, and the Case for A100 Training Infrastructure

The AS-4124GO-NART+ answers a procurement question that comes up in every AI infrastructure conversation: how do you get true 8-way NVLink training capability without committing frontier-tier budget? This is Supermicro's 4U platform built around the NVIDIA HGX A100 8-GPU baseboard, the same Delta-class GPU complex that powers the DGX A100, and it remains one of the most cost-effective routes to genuine distributed training infrastructure. The GPU complex and why NVSwitch is the whole point. Eight NVIDIA A100 80GB SXM4 GPUs mount directly to the HGX baseboard, interconnected through NVSwitch in an all-to-all mesh. Every GPU reaches every other GPU at 600 GB/s bidirectional NVLink bandwidth, roughly five times what PCIe Gen 4 delivers, without routing through the host. For distributed training, this is the difference between linear scaling and diminishing returns. Gradient synchronization after each training step is the bottleneck in multi-GPU work, and the NVSwitch fabric clears it. PCIe GPU servers, whatever their card count, cannot replicate this topology. If your workload is genuinely multi-GPU training rather than independent inference, the SXM-plus-NVSwitch architecture is what you are actually buying. Memory at node scale. Each A100 80GB carries HBM2e at approximately 2TB/s of bandwidth, and the node totals 640GB of GPU memory. In practice that supports full fine-tuning of models in the 30B to 70B class with model parallelism, training of smaller models at very large batch sizes, and inference serving of multiple 70B-class models simultaneously. Each A100 also partitions into up to 7 MIG instances, so a single node can present as many as 56 hardware-isolated GPU instances for multi-tenant inference, a density that makes this platform popular with GPU cloud and inference-serving operators. What NART+ means. Supermicro shipped this platform in two variants. The NART+ supports the higher-power A100 80GB SXM4 modules, which hold higher sustained clocks under continuous load. For a machine that will spend its life at full training utilization, sustained clocks translate directly into throughput, which is why the NART+ is the variant serious operators specify. The host platform. Dual Socket SP3 supports AMD EPYC 7002 and 7003 series processors up to 64 cores each, 128 cores and 256 threads per system. Thirty-two DIMM slots scale to 8TB of DDR4-3200 ECC memory. The working rule for an 8-GPU training node is system memory at or above total GPU memory, so configurations from 512GB to 1TB are the practical baseline here, with headroom far beyond. PCIe Gen 4 expansion supports up to one low-profile NIC per GPU, enabling GPUDirect RDMA with a 1:1 GPU-to-NIC ratio for InfiniBand or high-speed Ethernet clustering across multiple nodes. Hot-swap NVMe bays handle dataset storage, and redundant Titanium-efficiency power supplies carry the node's substantial draw with failover headroom. Where the A100 sits in the buying decision. The Hopper and Blackwell generations are faster per GPU, and nobody disputes it. The A100's case is different: it is the most production-proven AI accelerator ever deployed, every framework and serving stack is optimized for it, its CUDA ecosystem is fully mature, and the platform cost per unit of NVLink-connected training capability is dramatically below current-generation systems. For fine-tuning, mid-scale training, research clusters, university labs, and inference fleets, the economics of eight NVSwitch-connected A100s frequently beat a smaller count of newer GPUs. Teams already running A100 infrastructure also gain fleet consistency: same drivers, same containers, same operational playbook. Where it sits in MillionMiner's catalog. The SYS-741GE-TNRT covers flexible 1-to-4 PCIe GPU builds. This machine covers true 8-way NVSwitch training at Ampere economics. The HGX H100 and H200 platforms cover the tier above, and the DGX GB200 covers the frontier. Because GPU power class, CPU choice, memory, storage, and networking all change the build, every AS-4124GO-NART+ is configured to order: share your workload with MillionMiner, confirm the configuration with a specialist, and the system arrives tested, with customs handled, ready to rack.

Supermicro AS-4124GO-NART+: DGX A100-Class Training Infrastructure on Build-to-Order Economics

There are two ways to get eight A100 GPUs talking to each other at 600 GB/s. One is the NVIDIA DGX A100, a sealed appliance with a fixed configuration. The other is this server: the same NVIDIA HGX A100 8-GPU SXM4 baseboard with the same NVSwitch fabric, mounted in a Supermicro 4U platform where the CPUs, memory, storage, and networking are specified to your workload instead of decided for you. The GPU complex is the reason this machine exists. Eight A100 80GB SXM4 modules connect through NVSwitch in a full-mesh topology, so any GPU exchanges data with any other GPU at 600 GB/s without touching the CPU or PCIe bus. That is what makes distributed training scale: gradient synchronization across all eight GPUs completes fast enough that adding GPUs adds speed instead of overhead. The node carries 640GB of combined HBM2e at roughly 2TB/s per GPU, enough to train and fine-tune models in the tens of billions of parameters with data or model parallelism, or serve dozens of production inference workloads simultaneously. The NART+ designation matters: this variant supports the higher-power A100 80GB SXM4 modules, which sustain higher clocks under continuous training load than the standard-power version. The host side is dimensioned to feed eight hungry GPUs. Dual AMD EPYC 7002/7003 series processors deliver up to 128 cores and 256 threads, 32 DIMM slots scale to 8TB of DDR4-3200 ECC, and PCIe Gen 4 slots support one high-speed NIC per GPU for GPUDirect RDMA when you cluster multiple nodes. MillionMiner configures the full build against your training or inference plan and delivers worldwide DDP.

Need Help Choosing?

Our mining specialists can help you find the perfect miner for your setup and budget.

Supermicro AS-4124GO-NART+: 8x A100 80GB HGX Server

A genuine HGX A100 8-GPU server. Eight NVIDIA A100 80GB SXM4 GPUs sit on an NVSwitch baseboard, giving every GPU 600 GB/s NVLink to every other GPU and 640GB of combined HBM2e across the node. Dual AMD EPYC processors, up to 8TB of DDR4 ECC memory, PCIe Gen 4 throughout, and redundant Titanium power in a 4U chassis. This is the architecture behind the DGX A100, on a configurable platform: you choose CPUs, memory, storage, and networking, and MillionMiner builds, tests, and ships it worldwide DDP. The proven path to real distributed training for teams that want 8-way NVLink without Hopper-tier platform spend.

True HGX: 8x A100 on NVSwitch

Eight A100 80GB SXM4 GPUs in a full-mesh NVSwitch fabric. 600 GB/s GPU-to-GPU, 640GB pooled HBM2e. Real distributed training, not a PCIe compromise.

DGX A100 Capability, Configurable Build

The same 8-GPU Delta-class architecture as the DGX A100, with CPUs, memory, storage, and networking specified to your workload instead of fixed.

56 MIG Instances per Node

Each A100 splits into 7 isolated instances. One server presents up to 56 hardware-partitioned GPUs for multi-tenant inference and GPU cloud serving.

FAQ

Frequently Asked Questions

Yes. The AS-4124GO-NART+ is built around the NVIDIA HGX A100 8-GPU baseboard with SXM4 modules and NVSwitch interconnect, the same Delta-class GPU complex used in the DGX A100. This is distinct from PCIe GPU servers: every GPU connects to every other GPU at 600 GB/s through the NVSwitch mesh, which is the architecture distributed training actually requires.

Supermicro shipped this platform in two variants. The NART+ supports the higher-power A100 80GB SXM4 modules, which sustain higher clocks under continuous training load than the standard-power version. For a node running at full utilization around the clock, that sustained clock advantage converts directly into training throughput.

Economics and maturity. The A100 is the most production-proven AI accelerator ever deployed, with a fully mature CUDA ecosystem and every major framework optimized for it. The platform cost per unit of NVSwitch-connected training capability sits dramatically below Hopper-generation systems. For fine-tuning, mid-scale training, research clusters, and inference fleets, eight NVLinked A100s frequently outperform a smaller budget-equivalent count of newer GPUs. Teams extending existing A100 fleets also keep one operational playbook.

Same fundamental GPU architecture: eight A100 SXM4 GPUs on an HGX baseboard with NVSwitch. The DGX A100 is a sealed appliance with a fixed configuration and NVIDIA's bundled software and support. This Supermicro platform delivers the same GPU complex with a configurable host: you specify CPUs, memory capacity, storage, and networking against your actual workload, which typically lands at a meaningfully lower platform cost for the same training capability.

Full fine-tuning of models in the 30B to 70B class using model parallelism across the node. Training of smaller models at very large batch sizes. Inference serving of multiple 70B-class models simultaneously, or dozens of smaller models. With MIG, the node partitions into up to 56 isolated GPU instances for multi-tenant serving of 7B-to-13B-class models.

Dual AMD EPYC 7002 or 7003 series, up to 64 cores per socket, with 32 DIMM slots scaling to 8TB of DDR4-3200 ECC. The working rule for an 8-GPU training node is system memory at or above total GPU memory, so 512GB to 1TB is the practical baseline here. Data loading and preprocessing for eight GPUs needs real host throughput; an undersized host starves expensive accelerators. MillionMiner sizes this with you during configuration.

Yes. The PCIe Gen 4 expansion supports up to one low-profile NIC per GPU, enabling GPUDirect RDMA at a 1:1 GPU-to-NIC ratio. With InfiniBand or high-speed Ethernet adapters, multiple AS-4124GO-NART+ nodes form a training cluster where GPUs exchange data across nodes without CPU involvement. This is the standard scale-out path for workloads that outgrow a single 8-GPU node.

This is a dense 4U node with eight high-power SXM4 GPUs and dual server CPUs, carried by redundant Titanium-efficiency power supplies. Plan data-center-grade power delivery and high-volume front-to-back airflow; it is not an office-environment machine. MillionMiner confirms the exact power and thermal envelope of your specified configuration before the build, and hosting in MillionMiner's own facilities is available if you prefer not to provision power on-site.

Fully. The A100 runs current CUDA releases, and PyTorch, TensorFlow, JAX, TensorRT, Triton, vLLM, and every major training and serving stack maintain first-class Ampere support. A100 fleets still carry a substantial share of global cloud AI capacity, which keeps the ecosystem actively maintained. MIG, NVLink, and GPUDirect tooling are all mature on this platform.

The server is configured to order because GPU power class, CPUs, memory, storage, and networking all change the build. Submit your workload and deployment details through the quote form, confirm the configuration with a MillionMiner specialist, and the system is assembled, tested, and shipped worldwide DDP with duties and customs handled before delivery. Rack integration guidance and hosting in MillionMiner's data centers are both available.