Tencent Hunyuan (@TencentHunyuan) introduced WebCraftBench for the common failure mode where an agent claims a site is done while the homepage errors, buttons overlap, or an unrequested login flow appears. Agents operate the live app; coverage-guided exploration finds unreached paths; aesthetics, usability, and request match are scored. On 197 human-validated pairs it matches human preference 85.3%. Paper: https://arxiv.org/abs/2609.15387
Key Takeaways
- ✓Agents are evaluated on live apps, not only static screenshots or code diffs.
- ✓Coverage-guided exploration finds pages and flows that were never reached.
- ✓Matches human preference 85.3% on 197 validated pairs; paper arXiv:2609.15387.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.