Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
Not provided
Abstract
The paper introduces skill cascading attacks, a new threat paradigm in skill-based agent systems where malicious objectives are distributed across multiple skills, leading to harmful outcomes while appearing benign in isolation.
Reality Card
The introduction of skill cascading attacks reveals a significant gap in safety between individual skill integrity and overall system safety, necessitating new defenses that consider cross-skill interactions.
Development of SkillCascade, an automated multi-agent red-teaming framework, and SkillCascade-Bench, a benchmark of 213 validated cascading test cases.
The paper does not specify authors or detailed methodologies, which may hinder reproducibility.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.