vLLM claims support for NVIDIA Vera Rubin, reporting 7.8x throughput over GB200 in self-tests
The vLLM team announced support for NVIDIA's new Vera Rubin NVL72 platform, with blog-published self-test data showing 7.8x per-GPU throughput compared to GB200 NVL72 on AgentX workloads.
BedeutungWesentlichBeweisE2 nicht repliziertAufbereitungSchnell
The vLLM team has announced that its inference engine now supports NVIDIA's latest Vera Rubin NVL72 hardware platform.
According to preliminary test results published in the official blog, under AgentX workloads and maintaining matched interactivity standards, the per-GPU throughput on Vera Rubin NVL72 reached 7.8 times that of the previous generation GB200 NVL72. Additionally, in MLPerf vision-language model tests, it showed 3.7x higher throughput than GB300 NVL72.
This improvement is primarily driven by the Rubin platform's 5x NVFP4 FLOPS, 2.4x HBM bandwidth, and sixth-generation NVLink networking. vLLM introduced Rubin-tuned kernels via FlashInfer 0.7.0 and leveraged CUDA 13.4 locality domains to optimize memory access efficiency for Mixture-of-Experts (MoE) models.
It is important to note that these figures are internal benchmarks conducted by collaborators including vLLM, NVIDIA, and Red Hat, and have not yet been independently verified. Actual performance in production environments may vary depending on specific model architectures and workload types.