- Architecture
- NVIDIA Blackwell Ultra (dual-die, 208 billion transistors, same die design as B200), successor to Blackwell B200
- Memory
- 288 GB HBM3e (8 stacks, 12-high), 8 TB/s bandwidth per GPU
Computeper GPU, dense unless stated
- FP4 Tensor Core
- 15 PFLOPS dense (1.5x B200)
- FP8 Tensor Core
- 4.5 PFLOPS dense (9 PFLOPS with sparsity)
- FP16/BF16 Tensor Core
- 2.25 PFLOPS dense (4.5 PFLOPS with sparsity)
- TF32 Tensor Core
- 1.1 PFLOPS dense
- FP64
- About 1.2 TFLOPS, reduced versus B200; not positioned for FP64 HPC
- Attention layer (softmax / SFU)
- 2x B200 throughput
- Form Factor & TDP
- SXM module in HGX B300 baseboards, up to 1,400 W (configurable); liquid or air cooled per system design
Interconnects
- NVLink
- Fifth generation, 1.8 TB/s per GPU
- NVSwitch
- All-to-all bandwidth across 8 GPUs
- PCIe
- Gen 5 to the host
- HGX B300 Node
- 8 GPUs, 2.3 TB HBM3e per node
- Multi-Instance GPU
- Supported
- Engines
- Second-generation Transformer Engine (FP4 micro-tensor scaling), Decompression Engine, RAS Engine, Confidential Computing
- Software
- CUDA 12.x, NVIDIA AI Enterprise, NIM microservices, TensorRT-LLM, NVIDIA Dynamo, NeMo, PyTorch, JAX, vLLM, SGLang, Triton Inference Server