Supermicro
Model: AS-8126GS-NB3RT
Request Your Server Quote
Tell us your workload and deployment needs. Our specialist replies within 24 hours via email, WhatsApp, or call.
Why this server is quoted to order
These servers are configured and quoted to order. Your build, storage, networking, warranty, and rack integration determine the final price, and your delivery destination sets shipping and customs. Submit the form below and our specialist will reply within 24 hours with a full quote including hardware, warranty, and worldwide DDP delivery.
Thanks! Our specialist will reply within 24 hours with your custom quote.
How your server order works
Submit form
Share workload & deployment details
Get your quote
Complete pricing within 24 hours
Review with specialist
Finalise configuration and delivery
Delivery
Shipped and ready for deployment
Genuine
Tested hardware
Worldwide
Global shipping
Support
Mining experts
Supermicro's 8U A+ platform for the NVIDIA HGX B300 NVL8 baseboard: eight Blackwell Ultra B300 SXM GPUs with 288GB of HBM3e each, pooling 2.3TB of GPU memory per node behind fifth-generation NVLink and NVSwitch, with 800 Gb/s fabric ports integrated on the baseboard itself. Dual AMD EPYC 9005 or 9004 series processors, up to 6TB of DDR5 6400, eight E1.S NVMe bays, and six 6600W Titanium supplies carry it. The reasoning-era inference and training node, quoted and shipped worldwide DDP by MillionMiner.
Pricing, lead time, and hosting options. Personal advice from our sales team.
Pricing, lead time, and hosting options. Personal advice from our sales team.
Blackwell Ultra is the first GPU generation designed after the industry learned what reasoning models cost to serve. The B200 answered the training question; the B300 answers the inference-era question that followed it, and the changes are specific rather than cosmetic. This Supermicro A+ platform is the 8-GPU node those changes ship in. The GPU complex. Eight NVIDIA Blackwell Ultra B300 SXM GPUs mount to the HGX B300 NVL8 baseboard, each carrying 288GB of HBM3e, 60 percent more than the B200, pooling 2,304GB across the node. Fifth-generation NVLink joins every GPU at 1.8 TB/s through NVSwitch in full mesh. NVIDIA rates the B300's dense FP4 throughput at 1.5 times the B200, an uplift concentrated exactly where high-volume serving lives, and the second-generation Transformer Engine's micro-tensor scaling is what makes four-bit inference hold production accuracy. The memory bandwidth per GPU remains in the 8 TB/s class; the capacity is the headline, because capacity is what reasoning workloads exhaust first. Why capacity defines the reasoning era. A model that thinks before answering generates internal tokens, and every one of them occupies KV cache for the duration of the response. Long reasoning traces, extended context windows, and high-concurrency serving multiply that footprint together. On smaller nodes the result is offloading, batch-size collapse, or both, and serving margins go with them. A 2.3TB node holds frontier-scale weights and the caches they generate simultaneously, which is the difference between a model that fits and a deployment that profits. The integrated fabric, and what it changes about the chassis. The HGX B300 NVL8 baseboard carries its own networking: eight OSFP ports at 800 Gb/s, one per GPU, double the per-port rate of the prior generation's add-in adapters. This is why the spec sheet shows only two PCIe 5.0 expansion slots where B200-class systems carried a dozen; the twelve slots existed to hold NICs, and those NICs now live on the baseboard. For multi-node deployments the consequence is a cleaner build, a faster fabric, and GPUDirect RDMA topology guaranteed by design rather than assembled per order. The two remaining slots carry storage and management networking. The Supermicro host. Dual AMD EPYC 9005 or 9004 series processors, up to 192 cores per socket on the 9005 line, feed the preprocessing, tokenization, and data-loading work eight GPUs of this class demand. Twenty-four DDR5 modules at 6400 MT/s populate one per channel across twelve channels per processor and scale to 6TB, holding the full-bandwidth layout while sitting far above the GPU pool, the sizing rule this catalog teaches, cleared with room to spare. Eight front hot-swap E1.S NVMe bays keep datasets on-node, and six 6600W Titanium-level supplies in a redundant arrangement define a power envelope that is unambiguous about what this is: a data center machine at the top of the air-cooled single-node class. The decision against the HGX B200 in this catalog. The Gigabyte B200 node delivers 1.4TB and the same NVLink generation, and for training programs and serving fleets working within that memory envelope it remains the rational Blackwell buy. The B300 earns its position in three situations: reasoning and long-context inference fleets where KV cache depth decides batch size and therefore unit economics, deployments serving frontier-scale models that press past 1.4TB, and operators standardizing on the integrated 800G fabric for multi-node growth. Below those thresholds, the B200 wins on economics; at them, the B300 is the only node that fits. The decision against rack-scale. Above this machine sit the GB200 and GB300 class systems, where 72 GPUs share one NVLink domain under liquid cooling and facility engineering to match. The boundary is unchanged: if training parallelism genuinely needs more than eight GPUs in a single coherent domain, rack-scale is the answer. For everything that fits eight Blackwell Ultra GPUs, which now includes serving the largest deployed models, this node delivers the generation without the facility project and scales outward across its built-in 800G fabric. Export compliance and ordering. Blackwell-class accelerators are subject to US export controls, and MillionMiner confirms destination eligibility as part of every quote. Each system is configured to order, tested, and shipped worldwide DDP with duties and customs handled. Deployment planning and hosting in MillionMiner's own data centers are available for teams that prefer not to provision rack power at this scale on-site.
Inference economics changed when models started reasoning. A chain-of-thought response burns ten to one hundred times the tokens of a direct answer, every reasoning trace inflates the KV cache, and serving that profitably stopped being a compute problem and became a memory and precision problem. Blackwell Ultra is NVIDIA's direct answer, and this Supermicro platform is the 8-GPU form it ships in. The memory answer: 288GB of HBM3e per GPU, 60 percent more than the B200's 180GB, pooling 2.3TB across the node. That headroom is what holds frontier models alongside the long KV caches that reasoning traces and extended context windows generate, without the offloading that kills serving latency. The precision answer: NVIDIA rates the B300's dense FP4 at 1.5 times the B200, throughput aimed squarely at high-volume inference where four-bit serving with micro-tensor scaling holds accuracy and halves cost per token. The fabric answer: the HGX B300 NVL8 baseboard integrates eight 800 Gb/s OSFP ports directly, one per GPU at double the prior generation's rate, which is why this chassis needs only two PCIe expansion slots where B200 systems needed twelve. The cluster networking is no longer an add-on, it is part of the GPU complex. Supermicro's host keeps the proportions right. Dual AMD EPYC 9005 or 9004 series processors reach 192 cores per socket for preprocessing-heavy pipelines, twenty-four DDR5 modules at 6400 MT/s scale to 6TB, far above the 2.3TB GPU pool, and eight front hot-swap E1.S NVMe bays keep datasets local. Six 6600W Titanium supplies in a redundant arrangement carry the envelope. Quoted to configuration and shipped worldwide DDP by MillionMiner.
Our mining specialists can help you find the perfect miner for your setup and budget.
The largest single-node memory pool in this catalog: eight NVIDIA Blackwell Ultra B300 SXM GPUs at 288GB of HBM3e each, 2.3TB per node, joined by fifth-generation NVLink at 1.8 TB/s per GPU through NVSwitch. NVIDIA rates the B300's dense FP4 at 1.5 times the B200, and the HGX B300 NVL8 baseboard integrates its own 800 Gb/s fabric, eight OSFP ports, so clustering needs no add-in cards. Supermicro hosts it on dual AMD EPYC 9005 or 9004 processors with up to 6TB of DDR5 6400, eight E1.S NVMe bays, and six 6600W Titanium supplies in 8U. Quoted and shipped worldwide DDP by MillionMiner.
Eight Blackwell Ultra B300 SXM GPUs at 288GB each, 60 percent more than B200. Frontier weights and reasoning-scale KV caches fit together.
NVIDIA rates dense FP4 at 1.5x the B200. Chain-of-thought serving, long context, and high concurrency are what this node prices correctly.
Eight integrated OSFP ports at 800 Gb/s, one per GPU. Clustering is designed in, not added on, and the fabric runs at double the prior rate.
NVIDIA
Contact for price
Lenovo
Contact for price
Gigabyte
Contact for price
ASUS
Contact for price
HGX B300 NVL8 is NVIDIA's 8-GPU Blackwell Ultra building block: eight B300 SXM GPUs, the NVSwitch fabric joining them at 1.8 TB/s each, and integrated 800 Gb/s network ports, supplied to manufacturers like Supermicro. Against the B200, memory rises 60 percent to 288GB of HBM3e per GPU, NVIDIA rates dense FP4 throughput at 1.5 times higher, and the fabric ports double to 800 Gb/s while moving onto the baseboard itself.
2,304GB of pooled HBM3e across eight GPUs, fifth-generation NVLink at 1.8 TB/s per GPU through NVSwitch, dense FP4 throughput NVIDIA rates at 1.5 times the HGX B200, and eight integrated 800 Gb/s fabric ports. In workload terms: frontier-scale models serve with full KV cache headroom, models in the hundreds of billions of parameters fine-tune at full precision in one node, and reasoning fleets hold batch sizes that smaller memory pools collapse.
Because they multiply token counts. A model that reasons before answering generates internal chains of thought, every token of which occupies KV cache for the life of the response, and long contexts and high concurrency multiply the footprint further. When cache exhausts GPU memory, serving falls back to offloading or smaller batches, and unit economics fall with it. The B300's 288GB per GPU exists to hold weights and reasoning-scale caches simultaneously, and its dense FP4 uplift cuts the cost of every one of those extra tokens.
Three situations. Inference fleets serving reasoning or long-context workloads where KV cache depth sets batch size and margin. Models that press past the B200 node's 1.4TB. And operators standardizing on integrated 800G fabric for multi-node growth. Training programs and serving fleets comfortable inside 1.4TB are usually better served by the Gigabyte HGX B200 system, and MillionMiner models both in the quote.
Because the networking moved onto the GPU baseboard. B200-class systems carried a dozen slots largely to hold one 400 Gb/s adapter per GPU; the HGX B300 NVL8 integrates eight 800 Gb/s OSFP ports directly, so the one-port-per-GPU topology GPUDirect RDMA wants is built in at double the rate. The two remaining PCIe 5.0 slots carry storage and management networking, which is all that is left to add.
Dual AMD EPYC 9005 or 9004 series processors, reaching 192 cores per socket on the 9005 line, with twenty-four DDR5 modules at 6400 MT/s populated one per channel for full bandwidth and scaling to 6TB, far above the 2.3TB GPU pool. Storage runs eight front hot-swap E1.S NVMe bays for local datasets. The proportions follow the sizing rules the rest of this catalog teaches, with margin.
The boundary is the NVLink domain. This node joins eight GPUs in one coherent fabric on air-deployable infrastructure; rack-scale systems join 72 under liquid cooling and a facility engineering project. If training parallelism genuinely needs more than eight GPUs in one domain, rack-scale is the answer. For serving even the largest deployed models and for training that fits eight Blackwell Ultra GPUs, this node delivers the generation without the facility commitment and clusters outward over its integrated 800G fabric.
Over the baseboard's own fabric: eight 800 Gb/s OSFP ports, one per GPU, running GPUDirect RDMA so gradients and activations move between nodes without touching the CPU. The topology that B200-era systems assembled from add-in cards ships here by design. MillionMiner advises on switch and fabric architecture when a deployment grows past one machine.
Six 6600W Titanium-level redundant supplies define the envelope, which places this at the top of the air-deployable single-node class and unambiguously in data center territory. MillionMiner confirms the exact draw and thermal plan of your specified configuration during the quote, and hosting in MillionMiner's own facilities is available for teams that prefer not to provision rack power at this scale.
Submit your workload, scale, and deployment details through the quote form. A MillionMiner specialist confirms configuration, destination eligibility under the US export controls that apply to Blackwell-class accelerators, and the delivery plan. Every system is tested before shipment and delivered worldwide DDP with duties and customs handled. Rack integration guidance and hosted deployment are both available.