Supermicro

Supermicro GPU A+ Server AS-8126GS-NB3RT NVIDIA HGX B300 NVL8

Model: AS-8126GS-NB3RT

Request Your Server Quote

Tell us your workload and deployment needs. Our specialist replies within 24 hours via email, WhatsApp, or call.

Why this server is quoted to order

These servers are configured and quoted to order. Your build, storage, networking, warranty, and rack integration determine the final price, and your delivery destination sets shipping and customs. Submit the form below and our specialist will reply within 24 hours with a full quote including hardware, warranty, and worldwide DDP delivery.

How your server order works

1

Submit form

Share workload & deployment details

2

Get your quote

Complete pricing within 24 hours

3

Review with specialist

Finalise configuration and delivery

4

Delivery

Shipped and ready for deployment

Genuine

Tested hardware

Worldwide

Global shipping

Support

Mining experts

Supermicro's 8U A+ platform for the NVIDIA HGX B300 NVL8 baseboard: eight Blackwell Ultra B300 SXM GPUs with 288GB of HBM3e each, pooling 2.3TB of GPU memory per node behind fifth-generation NVLink and NVSwitch, with 800 Gb/s fabric ports integrated on the baseboard itself. Dual AMD EPYC 9005 or 9004 series processors, up to 6TB of DDR5 6400, eight E1.S NVMe bays, and six 6600W Titanium supplies carry it. The reasoning-era inference and training node, quoted and shipped worldwide DDP by MillionMiner.

Full Specifications

Model AS-8126GS-NB3RT
Form Factor 8U Rackmount
GPUs 8x NVIDIA B300 (Blackwell Ultra) SXM, 288GB HBM3e each
Total GPU Memory 2.3TB HBM3e (2,304GB pooled)
Interconnect NVIDIA NVLink 5 and NVSwitch, 1.8 TB/s per GPU
Processors Dual AMD EPYC™ 9005/9004 Series
Memory Support Up to 6TB DDR5, 6400 MT/s
DIMM Slots 24
Network Interfaces 8x OSFP 800 Gb/s (integrated on baseboard)
PCIe Slots 2x PCIe 5.0 x16 FHHL
Drive Bays 8x front hot-swap E1.S NVMe
Power Supplies 6x 6600W Redundant Titanium Level

Request a Bitcoin Miner Hosting Quote

Free quote, reply in 24h. No sales call.

4.4
star star star star star

4.7 / 5 on Trustpilot

Verified customer reviews

30,000+ miners delivered

Shipped worldwide since 2020

1,200+ customers globally

Trusted in 50+ countries

iso made-in-germany trustpilot
google-review

Get a Quote for the Supermicro GPU A+ Server AS-8126GS-NB3RT NVIDIA HGX B300 NVL8

Pricing, lead time, and hosting options. Personal advice from our sales team.

Reply within 24h via email, WhatsApp, or call.

Product Details

Supermicro AS-8126GS-NB3RT HGX B300 NVL8: Blackwell Ultra in Detail, the 2.3TB Node Mathematics, Integrated 800G Fabric, and the Decision Against B200 and Rack-Scale Systems

Blackwell Ultra is the first GPU generation designed after the industry learned what reasoning models cost to serve. The B200 answered the training question; the B300 answers the inference-era question that followed it, and the changes are specific rather than cosmetic. This Supermicro A+ platform is the 8-GPU node those changes ship in. The GPU complex. Eight NVIDIA Blackwell Ultra B300 SXM GPUs mount to the HGX B300 NVL8 baseboard, each carrying 288GB of HBM3e, 60 percent more than the B200, pooling 2,304GB across the node. Fifth-generation NVLink joins every GPU at 1.8 TB/s through NVSwitch in full mesh. NVIDIA rates the B300's dense FP4 throughput at 1.5 times the B200, an uplift concentrated exactly where high-volume serving lives, and the second-generation Transformer Engine's micro-tensor scaling is what makes four-bit inference hold production accuracy. The memory bandwidth per GPU remains in the 8 TB/s class; the capacity is the headline, because capacity is what reasoning workloads exhaust first. Why capacity defines the reasoning era. A model that thinks before answering generates internal tokens, and every one of them occupies KV cache for the duration of the response. Long reasoning traces, extended context windows, and high-concurrency serving multiply that footprint together. On smaller nodes the result is offloading, batch-size collapse, or both, and serving margins go with them. A 2.3TB node holds frontier-scale weights and the caches they generate simultaneously, which is the difference between a model that fits and a deployment that profits. The integrated fabric, and what it changes about the chassis. The HGX B300 NVL8 baseboard carries its own networking: eight OSFP ports at 800 Gb/s, one per GPU, double the per-port rate of the prior generation's add-in adapters. This is why the spec sheet shows only two PCIe 5.0 expansion slots where B200-class systems carried a dozen; the twelve slots existed to hold NICs, and those NICs now live on the baseboard. For multi-node deployments the consequence is a cleaner build, a faster fabric, and GPUDirect RDMA topology guaranteed by design rather than assembled per order. The two remaining slots carry storage and management networking. The Supermicro host. Dual AMD EPYC 9005 or 9004 series processors, up to 192 cores per socket on the 9005 line, feed the preprocessing, tokenization, and data-loading work eight GPUs of this class demand. Twenty-four DDR5 modules at 6400 MT/s populate one per channel across twelve channels per processor and scale to 6TB, holding the full-bandwidth layout while sitting far above the GPU pool, the sizing rule this catalog teaches, cleared with room to spare. Eight front hot-swap E1.S NVMe bays keep datasets on-node, and six 6600W Titanium-level supplies in a redundant arrangement define a power envelope that is unambiguous about what this is: a data center machine at the top of the air-cooled single-node class. The decision against the HGX B200 in this catalog. The Gigabyte B200 node delivers 1.4TB and the same NVLink generation, and for training programs and serving fleets working within that memory envelope it remains the rational Blackwell buy. The B300 earns its position in three situations: reasoning and long-context inference fleets where KV cache depth decides batch size and therefore unit economics, deployments serving frontier-scale models that press past 1.4TB, and operators standardizing on the integrated 800G fabric for multi-node growth. Below those thresholds, the B200 wins on economics; at them, the B300 is the only node that fits. The decision against rack-scale. Above this machine sit the GB200 and GB300 class systems, where 72 GPUs share one NVLink domain under liquid cooling and facility engineering to match. The boundary is unchanged: if training parallelism genuinely needs more than eight GPUs in a single coherent domain, rack-scale is the answer. For everything that fits eight Blackwell Ultra GPUs, which now includes serving the largest deployed models, this node delivers the generation without the facility project and scales outward across its built-in 800G fabric. Export compliance and ordering. Blackwell-class accelerators are subject to US export controls, and MillionMiner confirms destination eligibility as part of every quote. Each system is configured to order, tested, and shipped worldwide DDP with duties and customs handled. Deployment planning and hosting in MillionMiner's own data centers are available for teams that prefer not to provision rack power at this scale on-site.

Supermicro HGX B300 NVL8: The Node Built for Models That Think Before They Answer

Inference economics changed when models started reasoning. A chain-of-thought response burns ten to one hundred times the tokens of a direct answer, every reasoning trace inflates the KV cache, and serving that profitably stopped being a compute problem and became a memory and precision problem. Blackwell Ultra is NVIDIA's direct answer, and this Supermicro platform is the 8-GPU form it ships in. The memory answer: 288GB of HBM3e per GPU, 60 percent more than the B200's 180GB, pooling 2.3TB across the node. That headroom is what holds frontier models alongside the long KV caches that reasoning traces and extended context windows generate, without the offloading that kills serving latency. The precision answer: NVIDIA rates the B300's dense FP4 at 1.5 times the B200, throughput aimed squarely at high-volume inference where four-bit serving with micro-tensor scaling holds accuracy and halves cost per token. The fabric answer: the HGX B300 NVL8 baseboard integrates eight 800 Gb/s OSFP ports directly, one per GPU at double the prior generation's rate, which is why this chassis needs only two PCIe expansion slots where B200 systems needed twelve. The cluster networking is no longer an add-on, it is part of the GPU complex. Supermicro's host keeps the proportions right. Dual AMD EPYC 9005 or 9004 series processors reach 192 cores per socket for preprocessing-heavy pipelines, twenty-four DDR5 modules at 6400 MT/s scale to 6TB, far above the 2.3TB GPU pool, and eight front hot-swap E1.S NVMe bays keep datasets local. Six 6600W Titanium supplies in a redundant arrangement carry the envelope. Quoted to configuration and shipped worldwide DDP by MillionMiner.

Need Help Choosing?

Our mining specialists can help you find the perfect miner for your setup and budget.

Supermicro HGX B300 NVL8: 2.3TB of Blackwell Ultra

The largest single-node memory pool in this catalog: eight NVIDIA Blackwell Ultra B300 SXM GPUs at 288GB of HBM3e each, 2.3TB per node, joined by fifth-generation NVLink at 1.8 TB/s per GPU through NVSwitch. NVIDIA rates the B300's dense FP4 at 1.5 times the B200, and the HGX B300 NVL8 baseboard integrates its own 800 Gb/s fabric, eight OSFP ports, so clustering needs no add-in cards. Supermicro hosts it on dual AMD EPYC 9005 or 9004 processors with up to 6TB of DDR5 6400, eight E1.S NVMe bays, and six 6600W Titanium supplies in 8U. Quoted and shipped worldwide DDP by MillionMiner.

2.3TB of HBM3e in One Node

Eight Blackwell Ultra B300 SXM GPUs at 288GB each, 60 percent more than B200. Frontier weights and reasoning-scale KV caches fit together.

Built for the Reasoning Era

NVIDIA rates dense FP4 at 1.5x the B200. Chain-of-thought serving, long context, and high concurrency are what this node prices correctly.

800G Fabric on the Baseboard

Eight integrated OSFP ports at 800 Gb/s, one per GPU. Clustering is designed in, not added on, and the fabric runs at double the prior rate.

FAQ

Frequently Asked Questions

HGX B300 NVL8 is NVIDIA's 8-GPU Blackwell Ultra building block: eight B300 SXM GPUs, the NVSwitch fabric joining them at 1.8 TB/s each, and integrated 800 Gb/s network ports, supplied to manufacturers like Supermicro. Against the B200, memory rises 60 percent to 288GB of HBM3e per GPU, NVIDIA rates dense FP4 throughput at 1.5 times higher, and the fabric ports double to 800 Gb/s while moving onto the baseboard itself.

2,304GB of pooled HBM3e across eight GPUs, fifth-generation NVLink at 1.8 TB/s per GPU through NVSwitch, dense FP4 throughput NVIDIA rates at 1.5 times the HGX B200, and eight integrated 800 Gb/s fabric ports. In workload terms: frontier-scale models serve with full KV cache headroom, models in the hundreds of billions of parameters fine-tune at full precision in one node, and reasoning fleets hold batch sizes that smaller memory pools collapse.

Because they multiply token counts. A model that reasons before answering generates internal chains of thought, every token of which occupies KV cache for the life of the response, and long contexts and high concurrency multiply the footprint further. When cache exhausts GPU memory, serving falls back to offloading or smaller batches, and unit economics fall with it. The B300's 288GB per GPU exists to hold weights and reasoning-scale caches simultaneously, and its dense FP4 uplift cuts the cost of every one of those extra tokens.

Three situations. Inference fleets serving reasoning or long-context workloads where KV cache depth sets batch size and margin. Models that press past the B200 node's 1.4TB. And operators standardizing on integrated 800G fabric for multi-node growth. Training programs and serving fleets comfortable inside 1.4TB are usually better served by the Gigabyte HGX B200 system, and MillionMiner models both in the quote.

Because the networking moved onto the GPU baseboard. B200-class systems carried a dozen slots largely to hold one 400 Gb/s adapter per GPU; the HGX B300 NVL8 integrates eight 800 Gb/s OSFP ports directly, so the one-port-per-GPU topology GPUDirect RDMA wants is built in at double the rate. The two remaining PCIe 5.0 slots carry storage and management networking, which is all that is left to add.

Dual AMD EPYC 9005 or 9004 series processors, reaching 192 cores per socket on the 9005 line, with twenty-four DDR5 modules at 6400 MT/s populated one per channel for full bandwidth and scaling to 6TB, far above the 2.3TB GPU pool. Storage runs eight front hot-swap E1.S NVMe bays for local datasets. The proportions follow the sizing rules the rest of this catalog teaches, with margin.

The boundary is the NVLink domain. This node joins eight GPUs in one coherent fabric on air-deployable infrastructure; rack-scale systems join 72 under liquid cooling and a facility engineering project. If training parallelism genuinely needs more than eight GPUs in one domain, rack-scale is the answer. For serving even the largest deployed models and for training that fits eight Blackwell Ultra GPUs, this node delivers the generation without the facility commitment and clusters outward over its integrated 800G fabric.

Over the baseboard's own fabric: eight 800 Gb/s OSFP ports, one per GPU, running GPUDirect RDMA so gradients and activations move between nodes without touching the CPU. The topology that B200-era systems assembled from add-in cards ships here by design. MillionMiner advises on switch and fabric architecture when a deployment grows past one machine.

Six 6600W Titanium-level redundant supplies define the envelope, which places this at the top of the air-deployable single-node class and unambiguously in data center territory. MillionMiner confirms the exact draw and thermal plan of your specified configuration during the quote, and hosting in MillionMiner's own facilities is available for teams that prefer not to provision rack power at this scale.

Submit your workload, scale, and deployment details through the quote form. A MillionMiner specialist confirms configuration, destination eligibility under the US export controls that apply to Blackwell-class accelerators, and the delivery plan. Every system is tested before shipment and delivered worldwide DDP with duties and customs handled. Rack integration guidance and hosted deployment are both available.