While procedural skills enhance autonomous agents on complex tasks, generating verifiable training environments and instilling robust skill-invocation behaviors remains an open challenge. SkillGym introduces an end-to-end automated pipeline that filters reproducible offline skills, synthesizes 6.8k verifiable environments via a Builder-Reviewer architecture, and collects 19k verified successful trajectories. SFT fine-tuning boosts the relevant skill-invocation rate from 28% to 96% and enables Qwen3.5-9B to outperform a 397B parameter base model across two demanding skill benchmarks.

Key Takeaways

  • ✓Synthesizes 6.8k verifiable sandboxes and 19k verified trajectories, driving skill reading rates from 28% to 96%
  • ✓Enables fine-tuned Qwen3.5-9B to outperform an untrained 397B base model on two demanding skill benchmarks
  • ✓Demonstrates robust out-of-distribution generalization across entirely held-out skill repositories
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 While procedural skills are widely adopted across coding harnesses, base LLMs rarely invoke skills reliably on complex tasks—frequently ignoring relevant tools or hallucinating non-standard workflows, with baseline skill-invocation rates stagnating below 30%. Manually constructing diverse training sandboxes with executable verifiers is labor-prohibitive, bottlenecking agent post-training. ### 架构亮点与底层机制 SkillGym formalizes automated environment synthesis and skill post-training across four key pillars: 1. Reproducible Skill Ingestion: Crawls real-world skills, retaining deterministic workflows runnable offline. 2. Builder-Reviewer Environment Synthesis: Cooperatively generates difficulty-graded tasks paired with reference solutions and executable code verifiers, producing 6,800 isolated environments. 3. Verified Trajectory Harvesting: Executes and verifies 19,000 flawless end-to-end execution rollouts. 4. Scalable SFT Alignment: Standardizes supervised fine-tuning across 2B to 122B architectures. ### 权威 Benchmark 与实测跑分对比 Evaluated on four standard skill-use benchmarks across diverse model backbones: 1. Skill Invocation Surges from 28% to 96%: Post-training instills persistent tool exploration, lifting relevant skill inspection rates from 28% to 96%. 2. 9B Model Outperforms 397B Baseline: Fine-tuned Qwen3.5-9B outperforms a 397B untrained base model on two of four benchmarks. 3. Out-of-Distribution Robustness: Gains extend to held-out skills and minority task distributions, verifying general procedural competence over pattern memorization. ### 开发者实战落地与开箱指南 SkillGym is open-sourced on GitHub with environments, verified datasets, and SFT recipes. Engineering teams developing internal coding or DevOps agents can deploy SkillGym checkpoints or repurpose its Builder-Reviewer pipeline to synthesize training data for proprietary toolkits.