SafeActBench: NUS Investigates the Broken Evidence-to-Action Chain in Tool-Using Agents Across Six Operational Domains
Tool-using agents execute consequential modifications to external systems, yet nominal task success frequently conceals unverified actions executed without prior evidence. Researchers from the National University of Singapore (NUS) present a rigorous empirical study on the breakdowns along the evidence-to-action chain (arXiv:2610.07753). Evaluating ten model-harness configurations reveals a striking paradox: high static action assessment accuracy routinely masks brittle interactive execution. Failures overwhelmingly originate prior to execution: agents terminate prematurely before collecting requisite evidence, or initiate state-altering actions before foundational proofs are established. The authors formulate SafeActBench, spanning 656 rigorous cases across six operational domains and five protocols progressing from static judgment to complex multi-action workflows. Supported by a provenance-bound Evidence Ledger and deterministic trajectory evaluator, findings demonstrate that while single-step execution is stable once evidence is verified, multi-action workflows frequently collapse under unresolved prerequisites and truncated state transitions.
- •NUS presents SafeActBench across 656 operational cases and 6 domains to evaluate where the evidence-to-action chain breaks
- •Tests across 10 model-harness configurations show high static evaluation masks fragile execution, with over 70% of failures caused by premature actions
- •Introduces the Provenance-bound Evidence Ledger, proving multi-action workflows suffer 64.7% failure rates from unresolved prerequisites





