Vals AI launched Terminal-Bench 4.0: 66 new end-to-end terminal tasks (ship a service, prove a theorem, train a GPU kernel, write a forensic report, and more). Median expert effort is about 4 hours; grading is strict (full verifier suite or zero), and scores are avg@3. GPT-6 Astra won at 57.1%, 7.6pp ahead of Claude Fable 5.1 (49.5%) and 11.6pp ahead of Claude Opus 5 (45.5%). No other model cleared 30%; fourteen of 27 models scored zero on both hardware and media. Version 4.0 spans seven categories—software, science, ML, operations, hardware, security, and media—with roughly three-quarters of tasks outside traditional software work and no overlap with v2.1.
Key Takeaways
- ✓Terminal-Bench 4.0: 66 new terminal tasks, ~4h median expert effort, strict verifier + avg@3
- ✓GPT-6 Astra leads at 57.1% vs Fable 5.1 (49.5%) and Opus 5 (45.5%); no other model cleared 30%
- ✓Seven categories with ~3/4 tasks outside classic software engineering; no overlap with v2.1
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.