- Architecture
- AMD CDNA 3, chiplet design with 8 accelerator complex dies (XCD) on 4 I/O dies, 153 billion transistors
- Compute Units
- 304 compute units, 19,456 stream processors, 1,216 matrix cores
- Memory
- 192 GB HBM3, 5.3 TB/s bandwidth per GPU, 256 MB AMD Infinity Cache
Computeper GPU, peak
- FP64 Vector / FP64 Matrix
- 81.7 / 163.4 TFLOPS
- FP32 Vector
- 163.4 TFLOPS
- FP16/BF16
- 1,307.4 TFLOPS (2,614.9 TFLOPS with structured sparsity)
- FP8
- 2,614.9 TFLOPS (5,229.8 TFLOPS with structured sparsity)
- INT8
- 2,614.9 TOPS (5,229.8 TOPS with structured sparsity)
- Form Factor & TBP
- OAM module on the AMD Instinct MI300X platform baseboard, 750 W peak board power
Interconnects
- Infinity Fabric
- AMD Infinity Fabric links for 8-GPU all-to-all connectivity in the MI300X platform
- PCIe
- Gen 5 x16 to the host
- Platform
- AMD Instinct MI300X platform, 8 OAM GPUs and 1.5 TB of HBM3 per node
- Software
- ROCm 6, PyTorch, TensorFlow, JAX, vLLM, SGLang, Triton, Hugging Face, hipify for CUDA porting