- Architecture
- NVIDIA Blackwell (dual-die, 208 billion transistors, TSMC 4NP), successor to Hopper H100/H200
- Memory
- 180 GB HBM3e, 7.7 TB/s bandwidth per GPU
Computeper GPU, with sparsity
- FP4 Tensor Core
- 18 PFLOPS (9 PFLOPS dense)
- FP8/FP6 Tensor Core
- 9 PFLOPS (4.5 PFLOPS dense)
- FP16/BF16 Tensor Core
- 4.5 PFLOPS (2.25 PFLOPS dense)
- TF32 Tensor Core
- 2.2 PFLOPS (1.1 PFLOPS dense)
- FP32
- 75 TFLOPS
- FP64 / FP64 Tensor Core
- 37 / 40 TFLOPS
- Form Factor & TDP
- SXM module in HGX B200 baseboards, up to 1,000 W (configurable)
Interconnects
- NVLink
- Fifth generation, 1.8 TB/s per GPU
- NVSwitch
- All-to-all bandwidth across 8 GPUs
- PCIe
- Gen 5, 128 GB/s to the host
- Multi-Instance GPU
- Up to 7 MIG instances at 23 GB each
- Engines
- Second-generation Transformer Engine (FP4/FP6/FP8 micro-tensor scaling), Decompression Engine, RAS Engine, Confidential Computing with TEE-I/O
- Software
- CUDA 12.x, NVIDIA AI Enterprise, NIM microservices, TensorRT-LLM, NVIDIA Dynamo, NeMo, PyTorch, JAX, Triton Inference Server