On Oct 9 the vLLM team, with Inferact, Red Hat and NVIDIA, shipped early Vera Rubin NVL72 support: CUDA 13.4 nightly images already serve DeepSeek, Kimi, GLM and MiniMax. Vendor-reported results: up to 7.84x per-GPU throughput vs GB200 on SemiAnalysis AgentX (MiniMax M3, matched interactivity) and up to 3.7x vs GB300 NVL72 on the MLPerf Inference v6.1 VLM benchmark (Qwen3-VL-235B-A22B).

Key Takeaways

  • ✓Versus GB200 NVL72: 5x NVFP4 FLOPS, ~2.4x HBM4 bandwidth, 1.7x bidirectional NVLink, 2–4x faster softmax exponentials (vLLM blog)
  • ✓AgentX with MiniMax M3: up to 7.84x per-GPU throughput at matched interactivity, 5.18x under a 150 TPS cap (vendor-reported, early)
  • ✓MLPerf Inference v6.1: Qwen3-VL-235B-A22B on vLLM + Dynamo, up to 3.7x vs GB300 NVL72
  • ✓CUDA 13.4 locality domains split MoE weights for ~1.2x average MoE-layer speedup in small-token decode
  • ✓Image: vllm/vllm-openai:cu134-nightly (CUDA 13.4 + PyTorch 2.15)
vLLM adds NVIDIA Vera Rubin NVL72 support: up to 7.84x per-GPU throughput over GB200 on AgentX with MiniMax M3
🖼️Official Media
Click to view high-res
🧭

Not enough VRAM? Compare cloud API and self-hosting costs

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background

Agentic serving is bound by both throughput and interactivity; on GB200, MoE decode is limited by HBM weight reads and expert-parallel all-to-all traffic. Whether open-source engines run well on day one of NVIDIA's Vera Rubin matters to every deployer.

How it works

Per the vLLM blog, Rubin (sm107) is in the Blackwell family, so vLLM's sm100f kernels run unmodified. vLLM adds CUDA 13.4 locality domains that split MoE weights column-wise so each domain's SMs read only local HBM (~1.2x average MoE-layer speedup in small-token decode), Rubin-tuned FlashInfer 0.7.0 attention/GEMM/MoE kernels, and an FP8 MiniMax Sparse Attention prefill kernel.

Benchmarks (vendor-reported, early)

On SemiAnalysis AgentX with MiniMax M3: up to 7.84x per-GPU throughput vs GB200 at matched interactivity and 5.18x under a 150 TPS constraint. On MLPerf Inference v6.1 VLM (Qwen3-VL-235B-A22B, vLLM + Dynamo): up to 3.7x vs GB300 NVL72. Hardware: 5x NVFP4 FLOPS and ~2.4x HBM bandwidth vs GB200 NVL72.

Getting started

Pull vllm/vllm-openai:cu134-nightly (CUDA 13.4, PyTorch 2.15). Use --moe-backend flashinfer_cutedsl for NVFP4 MoE and --kv-cache-dtype fp8 with --attention-backend FLASHINFER for FP8 attention. Roadmap: MegaMoE, Kimi K3 MLA/KDA kernels, DeepSeek-V4.1-Flash Rubin kernels.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.