SkillPoison: Success Experiences Poison AI Skills
New research claims self-improving AI agents can be induced to develop harmful skills through seemingly normal successful experiences, with a 95.71% attack success rate.
ImportanciaMaterialEvidenciaE2 no replicadoAnálisisRápido
Self-improving LLM agents distill successful experiences into persistent skills, but this process carries a risk of covert manipulation.
Traditional attacks require injecting malicious triggers or false facts, which are easily detected and fail to accumulate. The SkillPoison framework demonstrates that even if all trajectories remain task-correct and pass verification, removing contextual conditions that constrain when behaviors apply can distort the skill extractor's generalization logic.
The team reports a 95.71% attack success rate across three benchmarks, noting that injected experiences remained correct and passed lexical inspection. Code is open-sourced, but these results are currently author-reported and have not yet been independently reproduced.