NVIDIA

NVIDIA DGX A100 Deep Learning System (8x A100, 640GB)

Model: DGX A100

Request Your Server Quote

Tell us your workload and deployment needs. Our specialist replies within 24 hours via email, WhatsApp, or call.

Why this server is quoted to order

These servers are configured and quoted to order. Your build, storage, networking, warranty, and rack integration determine the final price, and your delivery destination sets shipping and customs. Submit the form below and our specialist will reply within 24 hours with a full quote including hardware, warranty, and worldwide DDP delivery.

How your server order works

1

Submit form

Share workload & deployment details

2

Get your quote

Complete pricing within 24 hours

3

Review with specialist

Finalise configuration and delivery

4

Delivery

Shipped and ready for deployment

Genuine

Tested hardware

Worldwide

Global shipping

Support

Mining experts

The original NVIDIA-built AI appliance: eight A100 80GB SXM4 GPUs on six NVSwitches with 640GB of unified GPU memory, dual 64-core AMD EPYC CPUs, up to 2TB of system memory, and 30TB of Gen4 NVMe storage, delivered as one factory-integrated system running NVIDIA's full DGX software stack. Turnkey training infrastructure with single-vendor accountability, supplied and shipped worldwide DDP by MillionMiner.

Full Specifications

Model DGX A100
GPUs 8x NVIDIA A100
Total GPU Memory Up to 640 GB
NVLinks per GPU 12
Bi-directional Bandwidth 600 GB/s
NVSwitches 6x NVIDIA NVSwitch
AI Performance 5 petaFLOPS
MIG Up to 56 instances (7 per GPU)
System Memory Up to 2 TB
CPUs Dual 64-core AMD
Storage Up to 30 TB Gen4 NVMe SSD
Network Interface 10x Mellanox ConnectX-6
OS Ubuntu Linux

Request a Bitcoin Miner Hosting Quote

Free quote, reply in 24h. No sales call.

4.4
star star star star star

4.7 / 5 on Trustpilot

Verified customer reviews

30,000+ miners delivered

Shipped worldwide since 2020

1,200+ customers globally

Trusted in 50+ countries

iso made-in-germany trustpilot
google-review

Get a Quote for the NVIDIA DGX A100 Deep Learning System (8x A100, 640GB)

Pricing, lead time, and hosting options. Personal advice from our sales team.

Reply within 24h via email, WhatsApp, or call.

Product Details

NVIDIA DGX A100 Deep Learning System: Architecture, Software Stack, SuperPOD Scaling, and the Build-Versus-Buy Decision for 8-GPU AI Infrastructure

The DGX A100 occupies a specific place in AI infrastructure history: it is the machine that defined what an 8-GPU training node should look like. Every HGX A100 server on the market, including the Supermicro platforms in MillionMiner's own catalog, is built around the GPU baseboard NVIDIA designed for this system. Buying the DGX means buying the original, with everything NVIDIA layers on top of the silicon. The compute architecture. Eight NVIDIA A100 80GB SXM4 GPUs mount to the HGX baseboard and interconnect through six NVSwitches in a full-mesh topology. Each GPU carries 12 third-generation NVLinks, delivering 600 GB/s of bidirectional GPU-to-GPU bandwidth, with 4.8 TB/s of aggregate bisection bandwidth across the fabric. In practical terms, gradient synchronization across all eight GPUs completes fast enough that distributed training scales close to linearly, which is the entire reason SXM-plus-NVSwitch systems exist. The node totals 640GB of pooled HBM2e GPU memory and 5 petaFLOPS of AI performance, with INT8 inference throughput reaching 10 petaOPS. Multi-Instance GPU partitions each A100 into as many as seven isolated instances, so one DGX presents up to 56 hardware-partitioned GPUs for multi-tenant inference, a configuration that lets the same machine train by night and serve dozens of isolated workloads by day. The host platform is sized so the GPUs never wait. Dual 64-core AMD EPYC 7742 processors provide 128 cores and 256 threads for data loading, preprocessing, and augmentation. System memory reaches 2TB of DDR4. Storage splits between dedicated M.2 NVMe for the OS and 30TB of internal Gen4 NVMe for datasets, keeping training I/O off the boot path. Ten Mellanox ConnectX-6 interfaces include eight single-port 200Gb/s HDR InfiniBand adapters arranged one per GPU for GPUDirect RDMA, the topology that lets multiple DGX nodes exchange gradients directly between GPU memories across the network with no CPU in the path. The software stack, which is the half of the product white-box servers do not include. DGX OS, an Ubuntu-based distribution tuned for the hardware, ships installed. NVIDIA Base Command provides orchestration, scheduling, and fleet management. The NGC catalog supplies certified containers for every major training and inference framework, performance-validated against this exact system, so the gap between unboxing and first training run is measured in hours. NVIDIA AI Enterprise support covers the full stack under one contract: when a job fails, there is no integrator pointing at NVIDIA and NVIDIA pointing back. The scaling path is part of what you buy. The DGX A100 is the building block of NVIDIA's DGX SuperPOD reference architecture, the published blueprint for scaling from one node to clusters of twenty, one hundred, or more, with validated network topology, storage architecture, and management tooling. Organizations that start with one DGX inherit a documented road from single-node fine-tuning to cluster-scale training, rather than designing that road themselves. The build-versus-buy decision, stated honestly. MillionMiner sells both sides of it. The Supermicro AS-4124GO-NART+ in this catalog delivers the same eight-A100 NVSwitch GPU complex on a configurable host at a lower platform cost, and for budget-led teams comfortable owning their own integration and software stack, it is the rational pick. The DGX A100 is for the opposite buyer: enterprises and research institutions where procurement favors single-vendor accountability, where compliance requires a supported reference platform, and where the cost of engineers debugging integration issues exceeds the appliance premium. Time to first training run, one throat to choke, and a validated scaling path are the product. The DGX A100 remains the most widely deployed AI appliance ever built, with a software ecosystem that NVIDIA continues to maintain across current CUDA releases. As a 6U rackmount system drawing several kilowatts at full load, it belongs in a data center environment; MillionMiner confirms power and cooling requirements during configuration, ships worldwide DDP with customs handled, and offers hosting in its own facilities for teams that prefer not to provision infrastructure on-site.

NVIDIA DGX A100: Why Teams Pay for the Appliance Instead of Building the Server

Any competent integrator can bolt eight A100 GPUs into a chassis. What they cannot ship is the rest of the DGX A100: the system NVIDIA engineered around its own HGX baseboard, validated as one unit, and backed by one vendor for everything inside the box. The hardware is the reference design every 8-GPU A100 server imitates. Eight A100 80GB SXM4 GPUs interconnect through six NVSwitches, each GPU carrying 12 NVLinks for 600 GB/s of bidirectional bandwidth and 4.8 TB/s of aggregate fabric throughput, double the previous DGX generation. The node pools 640GB of GPU memory and delivers 5 petaFLOPS of AI performance. Dual 64-core AMD EPYC processors, up to 2TB of DDR4 system memory, 30TB of internal Gen4 NVMe, and ten Mellanox ConnectX-6 interfaces (eight at 200Gb/s for compute fabric, the rest for storage and management) round out a system with no component left for you to specify, source, or debug. The software is what the white-box alternatives cannot match. DGX OS arrives installed and tuned. NVIDIA Base Command handles orchestration, job scheduling, and cluster management. The NGC catalog supplies certified, performance-optimized containers for PyTorch, TensorFlow, RAPIDS, Triton, and hundreds of frameworks, tested specifically against DGX hardware. When something misbehaves anywhere in the stack, GPU, fabric, driver, or container, one support contract with one vendor owns the answer. That is the purchase logic. The DGX A100 is for teams whose cost of integration risk, debugging time, and delayed training runs exceeds the premium of the appliance. MillionMiner supplies the DGX A100 with worldwide DDP delivery and supports deployment planning, rack integration, and hosting.

Need Help Choosing?

Our mining specialists can help you find the perfect miner for your setup and budget.

NVIDIA DGX A100: The Turnkey 8-GPU AI Appliance

The DGX A100 is NVIDIA's own 8-GPU system, built, integrated, and supported end to end by the company that makes the GPUs. Eight A100 80GB SXM4 GPUs connect through six NVSwitches with 12 NVLinks per GPU at 600 GB/s, pooling 640GB of GPU memory behind 5 petaFLOPS of AI performance. Dual 64-core AMD EPYC CPUs, up to 2TB of system memory, 30TB of Gen4 NVMe storage, and ten Mellanox ConnectX-6 interfaces come pre-integrated, with DGX OS and NVIDIA's certified software stack installed before it ships. Power it on, pull a container from NGC, and start training the same day.

The Original 8x A100 Appliance

Eight A100 80GB GPUs, six NVSwitches, 640GB pooled memory, 5 petaFLOPS. The reference system every HGX A100 server is measured against.

Software Included, Risk Excluded

DGX OS, Base Command, and NGC certified containers arrive installed. One vendor supports the entire stack from GPU silicon to framework.

One Node Today, SuperPOD Tomorrow

Eight 200Gb/s ConnectX-6 fabric ports and NVIDIA's published SuperPOD architecture give you a validated path from one system to a training cluster.

FAQ

Frequently Asked Questions

The hardware GPU complex is similar by design, since other vendors build on NVIDIA's HGX baseboard. The difference is everything around it: NVIDIA engineers, validates, and supports the complete system as one product, with DGX OS, Base Command orchestration, and NGC certified containers preinstalled. White-box servers leave integration, OS tuning, and multi-vendor support coordination to you. The DGX removes those as risks.

Eight NVIDIA A100 80GB SXM4 GPUs with 640GB of total GPU memory, six NVSwitches with 12 NVLinks per GPU at 600 GB/s bidirectional, 5 petaFLOPS of AI performance, dual 64-core AMD EPYC CPUs (128 cores), up to 2TB of system memory, 30TB of internal Gen4 NVMe storage, and ten Mellanox ConnectX-6 network interfaces in a 6U rackmount chassis.

It lets any GPU exchange data with any other GPU at 600 GB/s without involving the CPU or PCIe bus, with 4.8 TB/s of aggregate bandwidth across the node. Distributed training synchronizes gradients between all GPUs after every step; on this fabric that synchronization is fast enough that eight GPUs deliver close to eight times the throughput of one. PCIe-connected GPU servers cannot replicate that scaling behavior.

Same class of GPU complex: eight A100 80GB SXM4 GPUs on an NVSwitch baseboard. The Supermicro is a configurable platform where you specify CPUs, memory, storage, and networking, typically at a lower platform cost, and you own the integration and software stack. The DGX is the sealed NVIDIA appliance with the full software stack and single-vendor support included. Budget-led teams pick the Supermicro; certainty-led teams pick the DGX. MillionMiner quotes both.

DGX OS, an Ubuntu-based distribution tuned for the hardware, comes installed. NVIDIA Base Command provides job scheduling, orchestration, and fleet management. The NGC catalog supplies certified, performance-optimized containers for PyTorch, TensorFlow, RAPIDS, Triton, and hundreds of other frameworks, validated against DGX hardware. The practical result is training workloads running within hours of power-on.

Full fine-tuning of models in the 30B to 70B parameter class with model parallelism across the node, training of smaller models at very large batch sizes, and inference serving of multiple 70B-class models simultaneously. With MIG, the system partitions into up to 56 isolated GPU instances, each suitable for serving models in the 7B to 13B class, which makes one DGX a complete multi-tenant inference platform.

Yes, and the system is built for it. Eight single-port 200Gb/s HDR InfiniBand ConnectX-6 adapters provide one fabric port per GPU for GPUDirect RDMA, letting GPUs in different nodes exchange data directly across the network. NVIDIA's DGX SuperPOD reference architecture documents the validated path from one node to clusters of twenty or more, including network topology, storage, and management design.

For a large share of real workloads, yes. The A100 is the most production-proven AI accelerator ever deployed, with mature support across current CUDA releases and every major framework. Newer generations are faster per GPU, but the DGX A100 delivers genuine NVSwitch-class distributed training at a platform cost well below current-generation appliances, which is exactly the trade fine-tuning teams, research institutions, and inference operators are looking for. The buying decision is workload economics, not generation chasing.

This is a 6U data center appliance weighing over one hundred kilograms and drawing several kilowatts at full training load, with high-volume front-to-back cooling. It is not an office machine. MillionMiner confirms the exact power, cooling, and rack requirements during configuration, and offers hosting in its own data centers for teams that prefer not to provision facilities.

Submit your workload and deployment details through the quote form. A MillionMiner specialist confirms configuration options, system memory, storage, support coverage, and deployment plan, then the system ships worldwide DDP with duties and customs handled before delivery. Rack integration guidance and hosted deployment in MillionMiner facilities are both available.