AMD ROCm Achieves 90-95% H100 Throughput for vLLM Inference
Tags Infrastructure · OSS · Enterprise

AMD's Advancing AI 2026 conference revealed ROCm achieving 90-95% of NVIDIA H100 throughput for standard LLM inference using PyTorch and vLLM, with significant cost advantages. However, TensorRT-LLM and FlashAttention 3 lack full ROCm equivalents, creating a dependency gap for teams using CUDA-specific libraries.
Technical significance
AMD's ROCm platform reaching 90-95% of H100 throughput represents a significant milestone in AI inference hardware competition. The cost advantages of AMD hardware combined with improved ROCm support for vLLM workloads could accelerate enterprise adoption of AMD infrastructure. However, the continued dependency on CUDA for TensorRT-LLM and FlashAttention 3 maintains NVIDIA's strategic advantage in specialized inference workloads.