SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
Yunhao Feng, Yifan Ding, Yingshui Tan, Boren Zheng, Xiaolong Li, Kun Zhai, Yishan Li, Yanming Guo, Wenke Huang
Abstract
Skill-based agent systems tackle complex tasks by composing reusable skills, improving modularity and scalability while introducing a largely unexamined security attack surface. We propose SkillTrojan, a backdoor attack that targets skill implementations rather than model parameters or training data. SkillTrojan embeds malicious logic inside otherwise plausible skills and leverages standard skill composition to reconstruct and execute an attacker-specified payload. The attack partitions an encrypted payload across multiple benign-looking skill invocations and activates only under a predefined trigger. SkillTrojan also supports automated synthesis of backdoored skills from arbitrary skill templates, enabling scalable propagation across skill-based agent ecosystems. To enable systematic evaluation, we release a dataset of 3,000+ curated backdoored skills spanning diverse skill patterns and trigger–payload configurations. We instantiate SkillTrojan in a representative code-based agent setting and evaluate both clean-task utility and attack success rate. Our results show that skill-level backdoors can be highly effective with minimal degradation of benign behavior, exposing a critical blind spot in current skill-based agent architectures and motivating defenses that explicitly reason about skill composition and execution. Concretely, on EHR SQL, SkillTrojan attains up to 97.2% ASR while maintaining 89.3% clean ACC on GPT-5.2-1211-Global. Code is available at https://github.com/Yunhao-Feng/SkillTrojan.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language ModelsBaolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong et al.ACL 2026 · 10 citations
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based AgentsHanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao et al.ICLR 2025
Related papers
- BadAgent: Inserting and Activating Backdoor Attacks in LLM AgentsYifei Wang, Dizhan Xue, Shengjie Zhang, Shengsheng QianACL 2024
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the WildYi Liu, Zhihao Chen, Yanjun Zhang, Gelei Deng et al.USENIX Security 2026 · 46 citations
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 197 citations
- Instruction Backdoor Attacks Against Customized LLMsRui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang et al.USENIX Security 2024 · 83 citations
- Backdooring Neural Code SearchWeisong Sun, Yuchen Chen, Guanhong Tao, Chunrong Fang et al.ACL 2023 · 18 citations
