Gigabyte
Model: G893-ZD1-AAX5
Request Your Server Quote
Tell us your workload and deployment needs. Our specialist replies within 24 hours via email, WhatsApp, or call.
Why this server is quoted to order
These servers are configured and quoted to order. Your build, storage, networking, warranty, and rack integration determine the final price, and your delivery destination sets shipping and customs. Submit the form below and our specialist will reply within 24 hours with a full quote including hardware, warranty, and worldwide DDP delivery.
Thanks! Our specialist will reply within 24 hours with your custom quote.
How your server order works
Submit form
Share workload & deployment details
Get your quote
Complete pricing within 24 hours
Review with specialist
Finalise configuration and delivery
Delivery
Shipped and ready for deployment
Genuine
Tested hardware
Worldwide
Global shipping
Support
Mining experts
Gigabyte's 8U platform for the NVIDIA HGX B200 baseboard: eight Blackwell B200 SXM GPUs with 180GB of HBM3e each, 1.4TB and 64 TB/s of memory per node, joined by fifth-generation NVLink at 1.8 TB/s per GPU through NVSwitch. Dual AMD EPYC 9005 or 9004 series processors, twenty-four DDR5 modules, eight Gen5 NVMe bays, and twelve 3000W Titanium supplies in a 6+6 redundant arrangement carry it, on air cooling. The current Blackwell frontier of single-node AI, quoted and shipped worldwide DDP by MillionMiner.
Pricing, lead time, and hosting options. Personal advice from our sales team.
Pricing, lead time, and hosting options. Personal advice from our sales team.
Every GPU generation gets sold as a revolution. The accurate way to evaluate Blackwell is to separate the physical changes from the marketing multipliers, and this server rewards that exercise, because the physical changes are substantial. The GPU itself. Each NVIDIA B200 is a dual-die package: two dies, each at the reticle limit of manufacturing, joined by a 10 TB/s die-to-die interface and presenting to software as a single GPU with 208 billion transistors. Each carries 180GB of HBM3e at roughly 8 TB/s of bandwidth, against 141GB at 4.8 TB/s on the H200. The second-generation Transformer Engine extends the precision ladder down to FP4 with micro-tensor scaling, hardware that tracks quantization scale at fine granularity so four-bit inference holds accuracy that earlier naive quantization lost. Fifth-generation NVLink doubles the fabric to 1.8 TB/s per GPU. The node mathematics. Eight B200 SXM GPUs on the HGX baseboard pool 1,440GB of HBM3e, 64 TB/s of aggregate memory bandwidth, and 144 petaFLOPS of FP4 compute, with NVSwitch joining every GPU to every other at full NVLink rate, 14.4 TB/s of aggregate fabric. In workload terms: models in the low hundreds of billions of parameters fine-tune at full precision inside one node, trillion-parameter-class models serve in real time across one node at FP4, long-context inference runs with KV cache headroom no Hopper node matches, and Mixture-of-Experts architectures route between experts on a fabric twice as fast as the H200 generation. NVIDIA's published comparisons, up to 15x real-time trillion-parameter inference and roughly 3x training against the H100 generation, are vendor figures, but the hardware deltas behind them are not. The Gigabyte platform. The host is dual AMD EPYC 9005 or 9004 series, reaching 192 cores per socket on the 9005 line, which matters for the tokenization, augmentation, and data-loading work that feeds eight GPUs of this class. Twenty-four DDR5 modules populate one per channel across twelve channels per processor, the configuration that sustains full memory bandwidth rather than trading it for capacity. Storage runs eight hot-swap Gen5 NVMe bays for local datasets. The PCIe layout is built for clustering: eight single-slot positions take one 400 Gb/s adapter per GPU, the one-to-one topology that GPUDirect RDMA wants, on NVIDIA Quantum-2 InfiniBand or Spectrum-X Ethernet, with four additional dual-slot positions for storage and management networking. Power is twelve 3000W 80 PLUS Titanium supplies in a 6+6 redundant arrangement, and the thermal design moves the full GPU complex on air, which keeps deployment inside ordinary data centers rather than the minority fitted for liquid. The decision against the H200 systems in this catalog. The Lenovo and ASUS HGX H200 platforms deliver 1.1TB per node on a 900 GB/s fabric, and for teams fine-tuning 70B-class models or serving within that memory envelope, they remain the rational buy. The B200 node earns its premium in three situations: inference fleets serving frontier-scale models where FP4 doubles tokens per watt, training runs where the doubled fabric and 60 percent bandwidth gain compress timelines that have business value, and workloads already pressing the H200's memory ceiling. Below those thresholds, Hopper economics win; at them, Blackwell does. The decision against rack-scale Blackwell. Above this machine sits the NVIDIA DGX GB200 class, where 72 GPUs share one NVLink domain at rack scale. The boundary is the NVLink domain itself: if your training parallelism needs more than eight GPUs in a single coherent fabric, rack-scale is the answer, and it brings liquid cooling, facility engineering, and a different order of commitment. For everything that fits eight Blackwell GPUs, which includes the large majority of enterprise training and nearly all inference serving, this node delivers the same generation without the facility project. Export compliance and ordering. Blackwell-class accelerators are subject to US export controls, and MillionMiner confirms destination eligibility as part of every quote. Each system is configured to order, tested, and shipped worldwide DDP with duties and customs handled. Deployment planning and hosting in MillionMiner's own data centers are available for teams that prefer not to provision rack power at this scale on-site.
Generational marketing is noisy, so here is the Blackwell step stated as numbers. Per GPU, memory rises from the H200's 141GB to 180GB of HBM3e, and bandwidth from 4.8 to roughly 8 TB/s, a 60 percent gain. The NVLink fabric doubles, 1.8 TB/s per GPU against 900 GB/s, through NVSwitch in full mesh. Per node, that compounds to 1.4TB of pooled GPU memory, 64 TB/s of aggregate bandwidth, and 144 petaFLOPS of FP4 compute. Each B200 is itself a dual-die design, two reticle-limit dies joined at 10 TB/s and presenting as one GPU with 208 billion transistors. The precision story matters as much as the bandwidth. Blackwell's second-generation Transformer Engine introduces FP4 with micro-tensor scaling, which is what turns trillion-parameter-class models from cluster problems into node problems for inference. NVIDIA's published figures put HGX B200 real-time inference at up to 15x the H100 generation on trillion-parameter workloads, with energy per token falling accordingly. For training, NVIDIA cites roughly 3x the H100 generation. Those are vendor benchmarks and should be read as such, but the architectural changes behind them, FP4, doubled fabric, 60 percent more memory bandwidth, are physical facts. Gigabyte's host platform keeps pace. Dual AMD EPYC 9005 or 9004 series processors reach 192 cores per socket for preprocessing-heavy pipelines, with twenty-four DDR5 modules populating one module per channel across twelve channels per processor, the layout that holds full memory bandwidth. Eight Gen5 NVMe bays keep datasets local, eight single-slot PCIe positions take one 400 Gb/s adapter per GPU for Quantum-2 InfiniBand or Spectrum-X clustering, and twelve 3000W Titanium supplies in 6+6 redundancy carry the envelope, on air cooling. Quoted to configuration and shipped worldwide DDP by MillionMiner.
Our mining specialists can help you find the perfect miner for your setup and budget.
This is the Blackwell tier: eight NVIDIA B200 SXM GPUs on the HGX baseboard, each carrying 180GB of HBM3e, pooling 1.4TB of GPU memory and 64 TB/s of bandwidth in one node. Fifth-generation NVLink joins every GPU at 1.8 TB/s through NVSwitch, double the H200 fabric, and the second-generation Transformer Engine adds FP4 precision that NVIDIA cites at up to 15x real-time inference for trillion-parameter models against the H100 generation. Gigabyte hosts it on dual AMD EPYC 9005 or 9004 processors with twenty-four DDR5 modules, eight Gen5 NVMe bays, and twelve 3000W Titanium supplies, cooled on air. Quoted and shipped worldwide DDP by MillionMiner.
Eight B200 SXM GPUs at 180GB each on NVSwitch. Trillion-parameter-class inference becomes a single-node, air-cooled purchase.
Second-generation Transformer Engine with micro-tensor scaling. NVIDIA cites up to 15x real-time inference versus the H100 generation.
Fifth-generation NVLink at 1.8 TB/s per GPU, 14.4 TB/s across the node. Gradient sync and MoE routing at twice Hopper bandwidth.
NVIDIA
Contact for price
Supermicro
Contact for price
Lenovo
Contact for price
Gigabyte
Contact for price
HGX B200 is NVIDIA's 8-GPU Blackwell building block: eight B200 SXM GPUs and the NVSwitch fabric joining them, supplied to manufacturers like Gigabyte who engineer complete servers around it. The physical changes against Hopper: each B200 is a dual-die package with 208 billion transistors, memory rises to 180GB of HBM3e per GPU at roughly 8 TB/s, NVLink doubles to 1.8 TB/s per GPU, and the second-generation Transformer Engine adds FP4 precision with micro-tensor scaling.
1,440GB of pooled HBM3e, 64 TB/s of aggregate memory bandwidth, 144 petaFLOPS of FP4 compute, and 14.4 TB/s of NVLink fabric through NVSwitch. In workload terms: full-precision fine-tuning of models in the low hundreds of billions of parameters, real-time FP4 serving of trillion-parameter-class models, and long-context inference with KV cache headroom beyond any Hopper node.
Three situations. Inference fleets at frontier scale, where FP4 roughly doubles tokens per watt and NVIDIA cites up to 15x real-time trillion-parameter inference against the H100 generation. Training programs where the doubled fabric and 60 percent bandwidth gain compress timelines with business value. And workloads already pressing the H200's 141GB per-GPU ceiling. Teams fine-tuning 70B-class models or serving within the Hopper envelope are usually better served by the Lenovo or ASUS HGX H200 systems, and MillionMiner will model both in the quote.
Usable, with engineering. The second-generation Transformer Engine implements micro-tensor scaling, tracking quantization scale at fine granularity so four-bit weights hold accuracy that naive quantization loses. Production serving stacks including TensorRT-LLM support it, and inference of large models is where it pays: memory footprint halves against FP8 and throughput roughly doubles. Training still runs FP8 and BF16; FP4 is an inference economics lever, and a large one at fleet scale.
Yes, that is the engineering premise of this Gigabyte platform: an 8U chassis with the airflow volume to hold the full HGX B200 complex at specification without direct liquid cooling. The practical consequence is deployment freedom, since no coolant distribution, plumbing, or liquid-loop maintenance is required and any data center with adequate power and conventional cooling qualifies. MillionMiner confirms the airflow and inlet temperature requirements for your site during the quote.
Core density and memory bandwidth. The EPYC 9005 series reaches 192 cores per socket, which feeds preprocessing-heavy pipelines for eight GPUs of this class, and the platform runs twelve DDR5 channels per processor with the twenty-four modules populated one per channel, the layout that sustains full bandwidth. The PCIe Gen5 lane budget also carries the one-adapter-per-GPU networking topology without compromise. The 9004 series remains available for teams standardized on it.
The boundary is the NVLink domain. This node joins eight GPUs in one coherent fabric; the GB200 class joins 72 at rack scale, with liquid cooling and facility engineering to match. If your training parallelism genuinely needs more than eight GPUs in a single domain, rack-scale is the answer. For everything that fits eight Blackwell GPUs, which covers most enterprise training and nearly all inference serving, this node delivers the same generation without the facility project, and scales outward over InfiniBand instead.
Through the one-adapter-per-GPU topology the PCIe layout is built for: eight single-slot positions take 400 Gb/s adapters on NVIDIA Quantum-2 InfiniBand or Spectrum-X Ethernet, giving GPUDirect RDMA a dedicated fabric port per GPU so gradients move between nodes without touching the CPU. Four additional dual-slot positions carry storage and management networking. MillionMiner advises on switch and fabric design when a deployment grows past one machine.
Twelve 3000W 80 PLUS Titanium supplies in a 6+6 redundant arrangement define the envelope, with the GPU complex alone capable of drawing eight kilowatts under sustained load before the host is counted. This is squarely a data center machine. MillionMiner confirms the exact draw of your specified configuration during the quote, and hosting in MillionMiner's own facilities is available for teams that prefer not to provision rack power at this scale.
Submit your workload, scale, and deployment details through the quote form. A MillionMiner specialist confirms configuration, destination eligibility under the US export controls that apply to Blackwell-class accelerators, and the delivery plan. Every system is tested before shipment and delivered worldwide DDP with duties and customs handled. Rack integration guidance and hosted deployment are both available.