Developer 3s Key Decision Metrics
On 2026-10-01 NVIDIA published DOCA AI agent skills on NVIDIA/skills: SKILL.md packs with real API signatures, hardware capability checks, and build constraints for Flow, GPUNetIO, PCC, RDMA, and more. Vendor 65-prompt eval: 19% checklist items without skills vs 100% with; RDMA demo used 73% less handwritten code and 46% fewer hardware commands. Installable into Claude Code, Codex, and other coding agents.
Key Takeaways
- ✓Blog 2026-10-01 + https://github.com/NVIDIA/skills
- ✓SKILL.md: real APIs, hardware capability manifests, build/preflight contracts
- ✓Vendor 65-prompt eval: 19% → 100% checklist; top failures without skills documented in blog
- ✓RDMA demo: −73% handwritten code, −46% hardware commands (vendor)
- ✓Install into Claude Code/Codex-class agents; keep human gates on firmware writes
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Core Background & Industry Pain Points
General coding agents invent plausible DOCA APIs/flags, skip hardware capability checks, and fail at link time—or worse, touch firmware without preflight. NVIDIA’s 2026-10-01 post frames the gap: without a machine-readable contract, agents guess from general training data.
Architecture Highlights & Internals
DOCA AI agent skills ship as SKILL.md packs with real signatures, capability manifests, build constraints, and failure mitigations across Flow, GPUNetIO, PCC, RDMA, and more. Firmware-class changes require maintenance windows, OOB assumptions, rollback, and cold power-cycle notes. Repo: NVIDIA/skills; install into Claude Code, Codex, and other Agent Skills hosts.
Authoritative Benchmarks & Measured Scores
Vendor-only numbers from the blog—do not invent external benches. 65 real prompts graded on required-answer checklists: 19% → 100% with skills; without-skills failure counts include API/flag misuse 59/65 and missing hardware checks 46/65. RDMA demo on BlueField-3: 189 vs 695 handwritten lines (−73%), 20 vs 37 hardware commands (−46%). Results depend on vendor prompts/graders.
Developer Hands-on Guide
Pull the relevant skills from NVIDIA/skills, install into your coding agent, verify device support before codegen, use pkg-config for link flags, and keep human change control for production DPU firmware writes.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.