A100

Values shown for A100 40 GB PCIe (the variant deployed in Galvani's a100-galvani partitions). * denotes "with sparsity" — without sparsity values are half.

Spec A100 40 GB PCIe A100 80 GB PCIe A100 40 GB SXM A100 80 GB SXM
Architecture NVIDIA Ampere NVIDIA Ampere NVIDIA Ampere NVIDIA Ampere
FP64 9.7 TFLOPS 9.7 TFLOPS 9.7 TFLOPS 9.7 TFLOPS
FP64 Tensor Core 19.5 TFLOPS 19.5 TFLOPS 19.5 TFLOPS 19.5 TFLOPS
FP32 19.5 TFLOPS 19.5 TFLOPS 19.5 TFLOPS 19.5 TFLOPS
TF32 Tensor Core 156 / 312* TFLOPS 156 / 312* TFLOPS 156 / 312* TFLOPS 156 / 312* TFLOPS
BFLOAT16 Tensor Core 312 / 624* TFLOPS 312 / 624* TFLOPS 312 / 624* TFLOPS 312 / 624* TFLOPS
FP16 Tensor Core 312 / 624* TFLOPS 312 / 624* TFLOPS 312 / 624* TFLOPS 312 / 624* TFLOPS
INT8 Tensor Core 624 / 1 248* TOPS 624 / 1 248* TOPS 624 / 1 248* TOPS 624 / 1 248* TOPS
GPU memory 40 GB HBM2 80 GB HBM2e 40 GB HBM2 80 GB HBM2e
Memory bandwidth 1 555 GB/s 1 935 GB/s 1 555 GB/s 2 039 GB/s
Max TDP 250 W 300 W 400 W 400 W
Multi-Instance GPU up to 7 MIGs @ 5 GB up to 7 MIGs @ 10 GB up to 7 MIGs @ 5 GB up to 7 MIGs @ 10 GB
Form factor PCIe PCIe SXM SXM
Interconnect NVLink Bridge for 2 GPUs (600 GB/s), PCIe Gen4 64 GB/s NVLink Bridge for 2 GPUs (600 GB/s), PCIe Gen4 64 GB/s NVLink 600 GB/s, PCIe Gen4 64 GB/s NVLink 600 GB/s, PCIe Gen4 64 GB/s

Architecture detail (from NVIDIA Ampere whitepaper): 6 912 FP32 CUDA cores, 3 456 FP64 CUDA cores, 432 third-generation Tensor cores per A100 GPU.


Last update: July 17, 2026
Created: July 17, 2026