Una tabella chiara per la velocità di inferenza AI, VRAM e throughput di training tra le GPU attuali — dai schede workstation all'hardware Blackwell dei data centre. Numeri di riferimento per aiutarla a scegliere la GPU giusta per il suo carico di lavoro.
Indice di inferenza AI più alto nel settore. Ogni numero di seguito è confrontato con l'RTX 3090 a 100.
Filtri per VRAM, ordini qualsiasi colonna e premi il + su due righe per inserirle direttamente nell'arena head-to-head sottostante.
| Seleziona | GPU ↕ | VRAM ↕ | Indice AI ▼ | FP16 TFLOPS ↕ | Formazione ↕ | Migliore per |
|---|---|---|---|---|---|---|
| 96 GB |
403
|
— | — | Largest LLMs without quantization, multi-tenant serving | ||
|
RTX Pro 5000 Blackwell
Flagship
|
48 GB |
285
|
— | — | High-throughput production inference | |
| 48 GB |
238
|
— | — | Best price-to-performance for serious inference | ||
| 32 GB |
207
|
419.1 | — | Fastest single-GPU inference & image generation | ||
| 40 GB |
152
|
312.0 | 1,396 img/s | Proven data-center training workhorse | ||
| 32 GB |
146
|
— | — | Balanced training + inference workstation | ||
|
RTX 4080 Super Pro
Mid
|
32 GB |
139
|
— | — | Mid-range inference & light fine-tuning | |
| 24 GB |
133
|
165.2 | 1,301 img/s | Best all-round price-to-performance card | ||
| 24 GB |
100
|
35.6 | 905 img/s | Budget-friendly entry into serious AI work | ||
|
V100 32GB
Entry
|
32 GB |
84
|
125.0 | — | Legacy data-center training | |
| 24 GB |
83
|
— | — | Compact workstation inference | ||
|
V100 16GB
Entry
|
16 GB |
62
|
125.0 | 833 img/s | Light training / dev environments | |
|
RTX 4070 Ti Super
Mid
|
16 GB |
56
|
— | — | Hobbyist projects & prototyping | |
|
RTX A4000
Entry
|
16 GB |
51
|
19.2 | — | Entry-level dev work, small models | |
| 96 GB | — | — | — | Power-efficient Max-Q variant for dense workstation builds | ||
| 96 GB | — | — | — | Boxed workstation edition, same silicon as Server Edition | ||
|
B200 SXM
Flagship
|
180 GB | — | 2,250.0 | — | Largest-scale multi-GPU training clusters | |
| 192 GB | — | 1,750.0 | — | Massive-context LLM training & serving | ||
|
Instinct MI350X
Flagship
|
288 GB | — | 2,306.9 | — | Massive VRAM pool for huge models, AMD ROCm stacks | |
|
Instinct MI325X
Flagship
|
256 GB | — | 1,307.4 | — | Large-model inference with huge memory headroom | |
| 141 GB | — | 989.5 | — | High-bandwidth-memory LLM serving at scale | ||
| 141 GB | — | 989.5 | — | PCIe/NVLink-bridge alternative to SXM for air-cooled racks | ||
|
Gaudi 3
Flagship
|
128 GB | — | 1,835.0 | — | Intel-based large-scale training alternative | |
|
Instinct MI300X
Flagship
|
192 GB | — | 1,307.4 | — | Single-GPU serving of very large open-weight models | |
| 80 GB | — | 989.5 | — | Industry-standard large-model training & inference | ||
|
RTX 6000 Ada
Flagship
|
48 GB | — | — | — | Top-tier workstation for training & rendering | |
| 32 GB | — | 419.1 | — | Blower-style RTX 5090 for dense multi-GPU builds | ||
| 24 GB | — | — | — | China-compliant RTX 5090 variant, cost-efficient rigs | ||
| 94 GB | — | 835.5 | — | Dual-GPU NVLink inference for very large models | ||
|
H100 PCIe
High-end
|
80 GB | — | 756.5 | — | PCIe-server LLM training & inference | |
| 80 GB | — | 756.5 | — | China-compliant H100 equivalent, reduced NVLink bandwidth | ||
| 80 GB | — | 312.0 | — | Export-compliant A100 alternative for restricted regions | ||
| 80 GB | — | 312.0 | — | Proven large-batch training workhorse | ||
| 40 GB | — | 312.0 | — | China-compliant A100 equivalent for large-scale training | ||
| 48 GB | — | 733.0 | — | Mixed training + inference + rendering server card | ||
| 32 GB | — | — | — | High-end workstation training & inference | ||
|
Instinct MI250X
High-end
|
128 GB | — | 383.0 | — | Large-memory AMD training clusters | |
|
RTX 4070 Ti
Mid
|
12 GB | — | 40.1 | — | Strong price-to-performance for local inference | |
|
L4
Mid
|
24 GB | — | 242.0 | — | Low-power inference & video AI at scale | |
| 48 GB | — | 181.1 | — | Graphics + AI mixed workloads on one card | ||
| 64 GB | — | 181.0 | — | PCIe AMD training/inference card | ||
|
Instinct MI100
Mid
|
32 GB | — | 184.6 | — | Earlier-gen AMD CDNA training/inference card | |
| 48 GB | — | 119.5 | — | Efficient server inference with large VRAM | ||
|
Radeon RX 7900 XTX
Mid
|
24 GB | — | 122.8 | — | AMD consumer flagship for local inference | |
|
RTX 4080 Super
Mid
|
16 GB | — | 104.4 | — | Strong mid-range inference & gaming/AI hybrid use | |
|
Radeon RX 7900 XT
Mid
|
20 GB | — | 103.2 | — | AMD mid-high consumer inference card | |
|
RTX 4080
Mid
|
16 GB | — | 97.5 | — | Solid mid-range inference workstation card | |
| 48 GB | — | 32.6 | — | Large-VRAM legacy workstation for bigger models | ||
|
A40
Mid
|
48 GB | — | 149.7 | — | Large-VRAM data-center inference card | |
| 24 GB | — | 165.0 | — | Efficient shared-server inference | ||
| 24 GB | — | 125.0 | — | Cost-efficient cloud inference card | ||
| 24 GB | — | 39.6 | — | Workstation training with large VRAM headroom | ||
|
RTX A6000
Mid
|
48 GB | — | 38.7 | — | Large-VRAM workstation for bigger models | |
|
RTX 3090 Ti
Mid
|
24 GB | — | 40.0 | — | High-VRAM consumer card for local models | |
|
RTX 3080 Ti
Mid
|
12 GB | — | 34.1 | — | Strong mid-range gaming/AI hybrid card | |
|
Arc A770
Mid
|
16 GB | — | 39.4 | — | Budget Intel card for OpenVINO/local inference | |
|
T4
Entry
|
16 GB | — | 65.0 | — | Very low-power cloud inference card | |
|
RTX A5000
Entry
|
24 GB | — | 27.8 | — | Reliable mid-VRAM workstation card | |
| 20 GB | — | 23.6 | — | Small-batch workstation training | ||
| 20 GB | — | 26.7 | — | Compact single-slot workstation inference | ||
| 20 GB | — | 38.4 | — | Small-form-factor workstation inference | ||
|
RTX 3080
Entry
|
10 GB | — | 29.8 | — | Affordable gaming/AI hybrid card | |
| 16 GB | — | 24.0 | — | Low-power small-form-factor inference | ||
| 16 GB | — | 22.3 | — | Legacy Turing workstation card, small models | ||
| 16 GB | — | 138.0 | — | Media/video-AI inference card | ||
|
RTX 3070
Entry
|
8 GB | — | 20.3 | — | Cheapest realistic entry point for small models | |
|
RTX 4060 Ti 16GB
Entry
|
16 GB | — | 22.1 | — | Budget 16GB card for local LLM experiments | |
| 12 GB | — | 8.0 | — | Ultra-low-power dev / edge inference | ||
|
RTX 4070
Entry
|
12 GB | — | 29.1 | — | Affordable current-gen dev/inference card | |
|
Radeon RX 7800 XT
Entry
|
16 GB | — | 74.4 | — | AMD mid-tier card with generous VRAM | |
| 16 GB | — | 18.7 | — | Legacy Pascal data-center training card | ||
|
RTX 4060
Entry
|
8 GB | — | 15.1 | — | Cheapest current-gen entry inference card | |
|
Instinct MI50
Entry
|
32 GB | — | 26.8 | — | Cheap secondhand AMD VRAM for local inference | |
|
Radeon RX 7600
Entry
|
8 GB | — | 43.5 | — | Budget AMD card for light local inference | |
| 8 GB | — | — | — | Ultra-low-power display/dev card, not AI-focused | ||
| 5 GB | — | — | — | Legacy budget workstation card, minimal AI use | ||
| 4 GB | — | 2.2 | — | Display-output card, not suitable for real AI workloads | ||
| 4 GB | — | — | — | Ultra-budget entry GPU, not recommended for AI workloads |
Mostrando 78 di 78 schede
Un server multi-GPU non viene valutato come una singola scheda — ciò che conta è la sua potenza di calcolo combinata di ogni GPU al suo interno. Il Total Compute qui sotto prende la cifra FP16 TFLOPS per scheda corrispondente dalla tabella sopra e la moltiplica per il numero di GPU, così che i server rimangano sulla stessa scala di benchmark. 48 sistemi attualmente nel nostro catalogo Hardware AI — toccate qualsiasi nome per le specifiche complete e il prezzo.
| Server | Configurazione GPU | VRAM totale | Potenza di calcolo totale (FP16 TFLOPS) | Prezzo |
|---|---|---|---|---|
|
NVIDIA DGX B200
DGX / Desktop
|
8x Blackwell GPUs, 1,440GB | 1,440 GB | 18,000.0 | $515,000 |
|
Gigabyte G893-SD1-AAX5
Rack Server
|
8x B200 SXM (HGX B200) | 1,440 GB | 18,000.0 | Richieda un preventivo |
|
Gigabyte G893-ZD1-AAX5
Rack Server
|
8x B200 SXM (HGX B200) | 1,440 GB | 18,000.0 | Richieda un preventivo |
|
HGX B200 8-GPU Baseboard
Baseboard
|
8x B200, 1,536GB HBM3e | 1,536 GB | 18,000.0 | Richieda un preventivo |
|
NVIDIA DGX H100
DGX / Desktop
|
8x H100 SXM5, 640GB | 640 GB | 7,916.0 | $6,700 |
|
NVIDIA DGX H200
DGX / Desktop
|
8x H200 SXM5, 1,128GB | 1,128 GB | 7,916.0 | Richieda un preventivo |
|
ASRock Rack 6U8X-EGS2
Rack Server
|
8x H200 SXM | 1,128 GB | 7,916.0 | $2,500 |
|
ASUS ESC N8-E11
Rack Server
|
8x H200 SXM (HGX) | 1,128 GB | 7,916.0 | $2,600 |
|
Dell PowerEdge XE9680
Rack Server
|
8x H100 SXM | 640 GB | 7,916.0 | $22,000 |
|
Gigabyte G593-SD2-AAX1
Rack Server
|
8x H100 SXM, 640GB | 640 GB | 7,916.0 | $5,600 |
|
Gigabyte G593-ZD2 (H100 Barebone)
Rack Server
|
8x H100 SXM5 (barebone) | 640 GB | 7,916.0 | $45,000 |
|
Gigabyte G593-ZD2 (H200 Barebone)
Rack Server
|
8x H200 SXM5 (barebone) | 1,128 GB | 7,916.0 | Richieda un preventivo |
|
Lenovo ThinkSystem SR675 V3
Rack Server
|
8x H100 SXM | 640 GB | 7,916.0 | $300,000 |
|
Lenovo HGX H200 8-GPU Server
Rack Server
|
8x H200 SXM, 141GB each | 1,128 GB | 7,916.0 | Richieda un preventivo |
|
Quanta S7PH H100
Rack Server
|
8x H100 SXM | 640 GB | 7,916.0 | $280,000 |
|
Quanta S7PH H200 (Barebone)
Rack Server
|
8x H200 SXM (barebone) | 1,128 GB | 7,916.0 | $20,000 |
|
HGX H100 Baseboard (Liquid)
Baseboard
|
8x H100 SXM5, 640GB, liquid-cooled | 640 GB | 7,916.0 | $3,400 |
|
HGX H100 8-GPU Baseboard
Baseboard
|
8x H100 SXM5, 640GB | 640 GB | 7,916.0 | $3,400 |
|
HGX H200 8-GPU Baseboard (Air)
Baseboard
|
8x H200 SXM, 1,128GB, air-cooled | 1,128 GB | 7,916.0 | Richieda un preventivo |
|
HGX H200 8-GPU Baseboard (Liquid)
Baseboard
|
8x H200 SXM, 1,128GB, liquid-cooled | 1,128 GB | 7,916.0 | Richieda un preventivo |
|
NVIDIA DGX H800
DGX / Desktop
|
8x H800 SXM5, 640GB | 640 GB | 6,052.0 | Richieda un preventivo |
|
Supermicro HGX H800 SYS-821GE
Rack Server
|
8x H800 SXM | 640 GB | 6,052.0 | $5,499 |
|
HGX H200 4-GPU Baseboard
Baseboard
|
4x H200 SXM, 564GB | 564 GB | 3,958.0 | $180,000 |
|
Lenovo HGX H200 4-GPU Board
Baseboard
|
4x H200 SXM, liquid-cooled | 564 GB | 3,958.0 | $7,800 |
|
Supermicro SYS-741GE-TNRT
Compact
|
4x H100 PCIe | 320 GB | 3,026.0 | $5,999 |
|
NVIDIA DGX A100
DGX / Desktop
|
8x A100 SXM4, 640GB | 640 GB | 2,496.0 | Richieda un preventivo |
|
Exeton Quasar 640X
Rack Server
|
8x A100 SXM, 640GB NVLink | 640 GB | 2,496.0 | Richieda un preventivo |
|
Gigabyte G492-ZD0 (Used)
Rack Server
|
8x A100 SXM | 640 GB | 2,496.0 | Richieda un preventivo |
|
Supermicro AS-4124GO-NART+
Rack Server
|
8x A100 HGX | 640 GB | 2,496.0 | $13,000 |
|
HGX A100 8-GPU Baseboard (640GB)
Baseboard
|
8x A100 SXM4, 640GB | 640 GB | 2,496.0 | $5,600 |
|
HGX A100 8-GPU Baseboard (320GB)
Baseboard
|
8x A100 SXM4 40GB, 320GB total | 320 GB | 2,496.0 | Richieda un preventivo |
|
NVIDIA DGX Station A100
DGX / Desktop
|
4x A100, 160GB | 160 GB | 1,248.0 | $85,000 |
|
RTX 5090 Full System
Compact
|
Pre-built RTX 5090 AI workstation | 32 GB | 419.1 | $6,000 |
|
ASUS Ascent GX10
DGX / Desktop
|
GB10 Grace Blackwell Superchip | 128 GB | — | Richieda un preventivo |
|
ASUS ESC8000A-E12P
Rack Server
|
Dual EPYC 9004, up to 8x GPU | — | — | Richieda un preventivo |
|
ASUS XA NB3I-E12
Rack Server
|
8x B300 NVL (HGX B300 NVL8) | — | — | $7,999 |
|
Gigabyte H263-S67 (2U 4-Node)
Compact
|
2U 4-node high-density chassis | — | — | $18,000 |
|
NVIDIA DGX Spark
DGX / Desktop
|
GB10 Grace Blackwell Superchip | 128 GB | — | Richieda un preventivo |
| Edge AI module, 64GB | 64 GB | — | Richieda un preventivo | |
|
QuantaGrid D74H-7U
Rack Server
|
8-GPU chassis (barebone) | — | — | Richieda un preventivo |
|
Supermicro AS-8125GS-TNHR (Refurb)
Rack Server
|
8-GPU server (refurbished) | — | — | Richieda un preventivo |
|
Supermicro AS-8126GS-NB3RT
Rack Server
|
8x B300 NVL (HGX B300 NVL8) | — | — | $450,000 |
|
Supermicro GH200 SuperServer
DGX / Desktop
|
NVIDIA GH200 Grace Hopper Superchip (used) | — | — | Richieda un preventivo |
|
Supermicro SYS-111C-NR-G1
Compact
|
1U GPU-ready server | — | — | $5,600 |
|
Supermicro SYS-212H-TN-G1
Compact
|
2U GPU-ready server | — | — | Richieda un preventivo |
|
Supermicro SYS-511E-WR-G1
Compact
|
1U GPU-ready server | — | — | $3,500 |
|
Supermicro SYS-521C-NR-G1
Compact
|
2U GPU-ready server | — | — | $4,500 |
|
Supermicro SYS-522GA-NRT
Compact
|
RTX PRO 6000 / L40S multi-GPU | — | — | Richieda un preventivo |
Hai già individuato un paio di preferiti sopra? Scegli due GPU e guardale confrontarsi su VRAM, indice di inferenza, calcolo FP16 e throughput di training — con una valutazione in linguaggio semplice su quale delle due vince e perché.
Iniziamo dai dati di riferimento pubblicamente disponibili sui benchmark GPU nel cloud — numeri verificabili che chiunque può controllare, senza supposizioni interne.

Ogni scheda viene misurata in ambito inferenza LLM, generazione di immagini e visione — il lavoro che l'hardware AI svolge effettivamente tutto il giorno.
Velocità di generazione dei token tra i modelli open-weight più diffusi, dagli assistenti leggeri da 8B ai pesi massimi da oltre 70B.
Elaborazione di modello di diffusione: throughput e latenza, coprendo tutto, dai pipeline turbo veloci alle rendering di qualità produttiva.
Comprensione multimodale delle immagini e throughput OCR di documenti sotto carico concorrente realistico.

Tutti i punteggi sono su una scala di 100 punti con RTX 3090 come riferimento — quindi due schede si confrontano immediatamente.
Una figura indipendentemente pubblicata per GPU singola (immagini/sec), mostrata quando esiste un benchmark corrispondente per quella scheda esatta.

Dati di riferimento solo a scopo di orientamento — i risultati reali variano in base alla stack software, ai driver, alla dimensione del batch e al carico di lavoro. Non si tratta di una garanzia di prestazioni.
Nuovo a questi termini? Leggi la guida in linguaggio semplice per scegliere una GPU AI
Se ha bisogno di una GPU specifica, di una configurazione completa di server o ha semplicemente una domanda riguardo ai numeri presentati in questa pagina — invii la sua richiesta e un vero ingegnere risponderà, di solito entro due ore.
Una guida approssimativa per abbinare il livello di GPU a ciò che si prevede di eseguire effettivamente.
Il massimo livello assoluto — progettato per aziende che forniscono AI a migliaia di utenti contemporaneamente.
Potenza di produzione seria per team esigenti e carichi di lavoro giornalieri intensi.
Il punto ideale — prestazioni elevate a un prezzo che un piccolo team può giustificare.
Un modo accessibile per entrare — imparare, creare prototipi e eseguire modelli più piccoli localmente.
È un numero relativo singolo che mostra come una GPU si comporta nell'inferenza LLM, nella generazione di immagini e nei carichi di lavoro di visione combinati, scalato in modo che l'RTX 3090 corrisponda sempre a 100. Un punteggio di 200 significa circa il doppio del throughput aggregato di un RTX 3090.
L'RTX Pro 6000 Blackwell attualmente guida la nostra classifica con un indice di inferenza AI di 403 — circa quattro volte la capacità complessiva di inference dell'RTX 3090 baseline — abbinato a 96 GB di VRAM per i più grandi LLM senza quantizzazione.
La RTX 5090 (AI Inference Index 207, 32 GB) e la RTX 4090 (index 133, 24 GB) offrono il miglior rapporto qualità-prezzo per inferenze su singola GPU e generazione di immagini. La variante RTX 4090 Pro aggiunge 48 GB di VRAM per modelli più grandi, mantenendo un rapporto qualità-prezzo da fascia consumer.
Per i più grandi LLM serviti senza quantizzazione, le schede di punta del data centre come l'RTX Pro 6000 Blackwell (96 GB), H200 (141 GB) e Instinct MI300X (192 GB) offrono spazio di memoria sufficiente per mantenere i modelli completi residenti. Per modelli da 8B a 30B, una singola RTX 5090 o RTX 4090 garantisce un ottimo throughput a una frazione del costo.
Sì — CPU GPU NVIDIA di ultima generazione e server AI completi sono disponibili nel nostro catalogo Hardware AI, con spedizione DDP in tutto il mondo e prezzi B2B su richiesta.
È una delle GPU in grado di AI più diffusamente distribuite, diventando un punto di riferimento pratico e ben consolidato per confrontare hardware più recenti e più vecchi.
No. La VRAM determina quali modelli e dimensioni di batch sono compatibili — non determina di per sé la velocità. Una scheda con meno VRAM può comunque ottenere un indice di inferenza più alto se la sua architettura e la larghezza di banda della memoria sono più performanti.
L'AMD Instinct MI350X guida con 288 GB, seguito dall'MI325X con 256 GB. Tra le schede NVIDIA, la B100 SXM offre 192 GB e l'H200 fornisce 141 GB di memoria ad alta larghezza di banda. Più VRAM consente di caricare modelli più grandi e batch più importanti senza dividerli su più GPU.
FP16 TFLOPS è una cifra teorica di picco di calcolo pubblicata dal produttore — un limite hardware grezzo. L'Indice di Inferenza AI è un punteggio misurato, basato sul carico di lavoro, costruito a partire da benchmark pubblicati su compiti LLM, immagini e visione. Una scheda può mostrare alte TFLOPS teoriche, ma un indice reale inferiore quando la larghezza di banda della memoria o il supporto software la limitano.
Entrambe sono GPU per data centre ad architettura Hopper valutate a 989,5 FP16 TFLOPS. La differenza è nella memoria: l'H200 possiede 141 GB di HBM3e più veloce rispetto agli 80 GB dell'H100, quindi supporta modelli più grandi e finestre di contesto più lunghe con una maggiore portata sostenuta. Per la maggior parte delle nuove distribuzioni di LLM, l'H200 rappresenta la scelta migliore.
Sì. Utilizzi l'Arena di confronto su questa pagina per selezionare due dei nostri 78 GPU elencati e visualizzate fianco a fianco il loro Indice di Inferenza AI, VRAM, TFLOPS FP16 e banda passante per l'addestramento, con il vincitore di ogni metrica evidenziato.
L'addestramento richiede molto più memoria e calcolo rispetto all'inferenza, quindi privilegia schede di fascia alta con molta VRAM e alte TFLOPS FP16 come la classe H100, H200, B200 o Instinct MI300X. La nostra colonna attraversoputs di Addestramento mostra le immagini per secondo su singola GPU dove esiste un benchmark pubblicato — ad esempio 1.396 img/s sulla A100 40GB contro 905 sulla RTX 3090.
Sì. Oltre a 78 GPU individuali,Elenco 48 server AI multi-GPU completi, inclusi sistemi di classe HGX e DGX basati su acceleratori H100, H200 e Blackwell, pronti per addestramenti su larga scala e inferenza di produzione. Contattateci per configurazioni e prezzi B2B.
Tutti e tre sono coperti. Accanto alla gamma completa di NVIDIA includiamo acceleratori AMD Instinct (MI350X, MI325X, MI300X, MI250X) e schede Radeon, oltre a Intel Gaudi 3 e Arc, così da poter confrontare tra i fornitori su un indice coerente.
Ogni dato è tratto da benchmark pubblicamente pubblicati e dai dati del produttore, aggregati in un indice coerente per orientarsi. I risultati nel mondo reale variano in base allo stack software, ai driver, alle dimensioni del batch e al carico di lavoro, quindi consideri i numeri come una guida comparativa piuttosto che una garanzia di prestazioni.
Revisioniamo e aggiorniamo le cifre man mano che vengono lanciate nuove generazioni di GPU e quando sono disponibili dati di benchmark più recenti.
Inviate la vostra domanda tramite il modulo di contatto sopra — o scriveteci su WhatsApp e riceverete una risposta in pochi minuti.