ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
Not provided
Abstract
The paper introduces ScopeBench, a benchmark for evaluating agents' adherence to engagement boundaries in security tasks, highlighting the importance of scope adherence in autonomous agents.
Reality Card
ScopeBench demonstrates that agents can be evaluated for both raw capability and scope adherence, revealing significant differences in performance across models.
The study found that scope adherence spans 34.4% to 86.7% across 8 models, with a judge identifying 331 violations missed by mechanical verification.
The main limitation is the potential for over-flagging by the agentic judge, which may affect the reliability of violation detection.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.