Developer 3s Key Decision Metrics
While procedural skills enhance autonomous agents on complex tasks, generating verifiable training environments and instilling robust skill-invocation behaviors remains an open challenge. SkillGym introduces an end-to-end automated pipeline that filters reproducible offline skills, synthesizes 6.8k verifiable environments via a Builder-Reviewer architecture, and collects 19k verified successful trajectories. SFT fine-tuning boosts the relevant skill-invocation rate from 28% to 96% and enables Qwen3.5-9B to outperform a 397B parameter base model across two demanding skill benchmarks.
Key Takeaways
- ✓Synthesizes 6.8k verifiable sandboxes and 19k verified trajectories, driving skill reading rates from 28% to 96%
- ✓Enables fine-tuned Qwen3.5-9B to outperform an untrained 397B base model on two demanding skill benchmarks
- ✓Demonstrates robust out-of-distribution generalization across entirely held-out skill repositories
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.