NVIDIA
Model: A100 40GB
Select Availability
Quantity
$9,500.00
Buying 10 or more? Get custom bulk pricing
Genuine
Tested hardware
Worldwide
Global shipping
Support
Mining experts
NVIDIA A100 40GB PCIe Tensor Core GPU. Ampere architecture (GA100, 7nm). 6,912 CUDA cores, 432 third-gen Tensor Cores, 40GB HBM2e on 5,120-bit bus at 1,555 GB/s bandwidth. 19.5 TFLOPS FP32, 156 TFLOPS TF32 (312 with sparsity), 312 TFLOPS FP16/BF16 (624 with sparsity). 250W TDP. Passive cooling dual-slot PCIe Gen 4 x16. MIG for up to 7 isolated GPU instances at 5GB each. NVLink bridge support for 2-GPU interconnect at 600 GB/s. "Original" denotes genuine NVIDIA/OEM production with enterprise warranty.
Pricing, lead time, and hosting options. Personal advice from our sales team.
Pricing, lead time, and hosting options. Personal advice from our sales team.
The A100 40GB occupies a specific position in MillionMiner's GPU catalog: it is the most battle-tested data center GPU available for buyers who prioritize ecosystem maturity and production reliability over bleeding-edge specifications. Five years after release, the A100 remains the workhorse of global AI cloud infrastructure and carries the deepest driver support, framework optimization, and enterprise deployment documentation of any GPU currently sold. Specifications tell the Ampere story. 6,912 CUDA cores and 432 third-gen Tensor Cores on the GA100 die at 7nm. 40GB HBM2e on a massive 5,120-bit bus delivering 1,555 GB/s memory bandwidth. HBM bandwidth is the A100's structural advantage over GDDR-based GPUs: the 1,555 GB/s figure beats most consumer and workstation GPUs on raw bandwidth per byte, which matters for memory-bound inference and HPC workloads. 19.5 TFLOPS FP32. TF32 Tensor performance at 156 TFLOPS (312 with sparsity) provides up to 20x throughput over the previous Volta generation for AI training with zero code changes. FP16/BF16 at 312 TFLOPS (624 with sparsity). Strong FP64 Tensor at 19.5 TFLOPS for scientific double-precision HPC. MIG (Multi-Instance GPU) creates up to 7 fully isolated instances at 5GB each with dedicated memory, cache, and compute. This is more granular than the RTX PRO 6000's 4-instance MIG, making the A100 better suited for multi-tenant inference serving where many small models run concurrently. NVLink bridge connects two A100 PCIe cards at 600 GB/s bidirectional, effectively creating a unified 80GB pool with high-speed interconnect. This is a capability neither the RTX PRO 6000 nor the RTX 5090 offers. The honest limitation is 40GB. In 2026, 40GB VRAM constrains large language model workloads to approximately 7B to 13B parameters at FP16 for fine-tuning, or approximately 25B at INT8 for inference. Models like Llama 3 70B at FP16 require 140GB and will not fit. For LoRA and QLoRA fine-tuning of 7B to 13B models, the 40GB is sufficient and cost-effective. For pure inference on quantized models (GPTQ, AWQ, GGUF at 4-bit), larger models fit because quantization compresses the memory footprint by 4x to 8x. Compared to MillionMiner's other GPU offerings. Against the A100 80GB Custom ($7,900 to $8,200): the 80GB doubles memory for similar price, making it the better buy for most AI workloads unless the "Original" genuine warranty matters over the "Custom" designation. Against the RTX PRO 6000 Workstation ($10,000 to $11,000): newer Blackwell architecture, 96GB GDDR7, 125 TFLOPS FP32 versus 19.5, but no NVLink and GDDR versus HBM bandwidth characteristics differ. For new single-GPU AI deployments, the RTX PRO 6000 wins on compute. For proven multi-GPU data center builds with NVLink and MIG-heavy workloads, the A100 ecosystem remains stronger. GPU mining: the A100 40GB can mine GPU-mineable algorithms but this is not the card's value proposition in 2026. Mining profitability is negative at standard electricity rates. The A100's 250W TDP and HBM memory offer some efficiency advantages over GDDR-based cards on memory-hard algorithms, but the economics do not justify the $8,000 acquisition cost for mining alone. Production status note: NVIDIA reportedly began winding down A100 manufacturing in 2024. Remaining new inventory is finite. Buyers building A100-based fleets should secure inventory while genuine new units are available, as the market will shift to refurbished and used inventory over time. 250W TDP passively cooled. Dual-slot PCIe Gen 4 x16. Brand New at MillionMiner $8,000.
The A100 40GB is not the newest GPU in MillionMiner's catalog and that is precisely the point. Released in 2020 on NVIDIA's Ampere architecture, the A100 has five years of production deployment behind it. Every major cloud provider runs A100 fleets (AWS P4 instances, GCP A2, Azure NC A100). Every major ML framework is optimized for it. Every enterprise support contract, driver update, and CUDA toolkit release is verified against it. No other GPU has this depth of proven production infrastructure. At $8,000 for a genuine Brand New unit, the A100 40GB sits at the competitive floor for verified new inventory. Gray-market units below $1,800 risk firmware issues, untested memory, and incompatibility with CUDA 12.3+. "Original" at MillionMiner means genuine NVIDIA/OEM production with full enterprise warranty. Core specifications: 6,912 CUDA cores, 432 third-gen Tensor Cores (supporting FP16, BF16, TF32, INT8, INT4, and FP64 acceleration), 40GB HBM2e at 1,555 GB/s on a 5,120-bit bus. 19.5 TFLOPS FP32, 156 TFLOPS TF32 (312 with sparsity), 312 TFLOPS FP16 (624 with sparsity). 250W TDP passively cooled in a dual-slot PCIe Gen 4 form factor. MIG divides one A100 into up to 7 isolated instances at 5GB each with guaranteed QoS. NVLink bridge connects two A100s at 600 GB/s for double memory and interconnect bandwidth. The 40GB VRAM handles inference on quantized models up to approximately 25B parameters and LoRA fine-tuning on 7B to 13B models. For larger models, the A100 80GB ($7,900 to $8,200 at MillionMiner) or RTX PRO 6000 96GB ($10,000+) provides more headroom.
Our mining specialists can help you find the perfect miner for your setup and budget.
NVIDIA's Ampere data center GPU. 6,912 CUDA cores, 432 third-gen Tensor Cores, 40GB HBM2e at 1,555 GB/s. 19.5 TFLOPS FP32, up to 624 TFLOPS FP16 with sparsity. 250W TDP passively cooled. MIG for up to 7 isolated instances. NVLink bridge for 2-GPU 600 GB/s interconnect. PCIe Gen 4 x16. The most widely deployed AI accelerator in cloud infrastructure worldwide. Proven ecosystem across 2,000+ applications. "Original" genuine NVIDIA production, Brand New at MillionMiner $8,000.
5,120-bit memory bus. HBM bandwidth advantage over GDDR-based GPUs. Proven across AWS, GCP, Azure A100 cloud fleets.
More granular multi-tenancy than newer GPUs. 7 isolated instances at 5GB each. NVLink connects 2 GPUs at 600 GB/s.
Lowest TDP in MillionMiner's professional GPU lineup. Passive cooling for server chassis. Five years of production proven stability.
Genuine NVIDIA or authorized OEM production A100 40GB PCIe with full enterprise warranty. Distinguishes this from the "A100 80G Custom" variant also in MillionMiner's catalog, which may use a modified or aftermarket configuration. "Original" means factory specifications, verified firmware, and standard NVIDIA warranty coverage. Brand New condition.
Ecosystem maturity. The A100 has the deepest driver support, ML framework optimization, and enterprise deployment documentation of any GPU. Every cloud provider runs A100 fleets. Every production inference pipeline is tested against it. If you need proven reliability over cutting-edge specs, the A100 remains the safest infrastructure investment. At $8,000, it is also the lowest-cost genuine NVIDIA data center GPU in MillionMiner's catalog.
Inference on quantized models up to approximately 25B parameters at INT8. LoRA/QLoRA fine-tuning on 7B to 13B models at FP16. Full fine-tuning on models up to approximately 7B at FP16 with optimizer offloading. Inference on smaller production models (BERT, ViT, stable diffusion, ResNet class) without constraints. Cannot handle 70B+ models at FP16; those need the 80GB variant or RTX PRO 6000 96GB.
Same GA100 chip, same core counts, same compute TFLOPS. The 80GB doubles memory (80GB HBM2e, 1,935 GB/s for 80GB PCIe versus 1,555 GB/s for 40GB) and increases MIG instance size (10GB per instance versus 5GB). 80GB TDP is 300W versus 250W. MillionMiner's A100 80GB Custom is priced at $7,900 to $8,200, comparable to the 40GB Original at $8,000. Unless the "Original" genuine warranty matters versus the "Custom" designation, the 80GB is the better value for most workloads.
PCIe A100 (this product) fits standard server and workstation motherboards via PCIe Gen 4 x16 slot. 250W TDP. NVLink via bridge between 2 GPUs. SXM A100 uses NVIDIA's HGX baseboard with direct NVLink across up to 8 GPUs via NVSwitch. 400W TDP. Higher performance but requires purpose-built HGX server infrastructure. The PCIe variant is the versatile option for existing server platforms.
Multi-Instance GPU creates up to 7 fully isolated instances on one A100, each with 5GB dedicated memory, dedicated cache, and compute resources with guaranteed QoS. Works with Kubernetes, containers, and hypervisor-based virtualization. The 7-instance granularity exceeds newer GPUs like the RTX PRO 6000 (4 instances), making the A100 better for multi-tenant inference serving with many small concurrent models.
Yes, via NVLink bridge connecting two PCIe A100 GPUs at 600 GB/s bidirectional bandwidth. This effectively creates a unified 80GB memory pool with high-speed interconnect. Neither the RTX PRO 6000 nor the RTX 5090 offers NVLink. The A100's NVLink capability is a genuine architectural differentiator for 2-GPU scaling on memory-bound workloads.
Technically yes on GPU-mineable algorithms, but this is not a mining card in 2026. Hashrate.no data shows GPU mining profitability is negative at standard electricity rates for data center GPUs at this price tier. The 250W TDP and HBM memory provide some efficiency on memory-hard algorithms, but the $8,000 acquisition cost makes mining ROI impractical. Buy for AI compute and HPC.
Yes. NVIDIA continues CUDA and driver support for Ampere architecture across current CUDA toolkit releases (12.x). No announced end-of-support date. Data center GPUs historically receive driver support for many years after production ends. Current ML frameworks (PyTorch, TensorFlow, JAX, TensorRT, Triton) all maintain A100 optimization.
A100 40GB: 19.5 TFLOPS FP32, 40GB HBM2e at 1,555 GB/s, NVLink, 7 MIG instances, 250W, $8,000. RTX PRO 6000: 125 TFLOPS FP32, 96GB GDDR7 at 1,792 GB/s, no NVLink, 4 MIG instances, 600W, $10,000+. The RTX PRO 6000 wins on raw FP32 compute (6.4x) and memory capacity (2.4x). The A100 wins on NVLink interconnect, MIG granularity, power efficiency per GB, and production ecosystem maturity. Different tools for different deployment priorities.