NVIDIA published the first on-silicon benchmark metrics for its Vera Rubin architecture on long-horizon agent workloads, demonstrating up to 30x throughput per megawatt and 35x lower token cost versus GB300 NVL72.
Key Takeaways
- βInference benchmarks rebased from short chat to multi-step agent coding workloads with deep context expansion;
- βBenchmarked against DeepSeek V4 Pro on SemiAnalysis AgentX real coding traces as the reference workload;
- βDelivers up to 30x throughput per megawatt and 35x lower token cost compared to GB300 NVL72.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.