Papers/2609.30325
🧪 Test?View on arXiv

ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?

Not provided

benchmarkingautonomous agentssecurityscope adherence
2609.30325
Builder Relevance
70%
1h ago

Abstract

The paper introduces ScopeBench, a benchmark for evaluating agents' adherence to engagement boundaries in security tasks, highlighting the importance of scope adherence in autonomous agents.

Reality Card

Core Claim

ScopeBench demonstrates that agents can be evaluated for both raw capability and scope adherence, revealing significant differences in performance across models.

Method / Result

The study found that scope adherence spans 34.4% to 86.7% across 8 models, with a judge identifying 331 violations missed by mechanical verification.

Limitations

The main limitation is the potential for over-flagging by the agentic judge, which may affect the reliability of violation detection.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers