BAM-ICL: Causal Hijacking In-Context Learning with Budgeted Adversarial Manipulation
Rui Chu, Bingyin Zhao, Hanling Jiang, Shuchin Aeron, Yingjie Lao
摘要
Recent research shows that large language models (LLMs) are vulnerable to hijacking attacks under the scenario of in-context learning (ICL) where LLMs demonstrate impressive capabilities in performing tasks by conditioning on a sequence of in-context examples (ICEs) (i.e., prompts with task-specific input-output pairs). Adversaries can manipulate the provided ICEs to steer the model toward attackerspecified outputs, effectively "hijacking" the model's decision-making process. Unlike traditional adversarial attacks targeting single inputs, hijacking attacks in LLMs aim to subtly manipulate the initial few examples to influence the model's behavior across a range of subsequent inputs, which requires distributed and stealthy perturbations. However, existing approaches overlook how to effectively allocate the perturbation budget across ICEs. We argue that fixed budgets miss the potential of dynamic reallocation to improve attack success while maintaining high stealthiness and text quality. In this paper, we propose BAM-ICL, a novel budgeted adversarial manipulation hijacking attack framework for in-context learning. We also consider a more practical yet stringent scenario where ICEs arrive sequentially and only the current ICE can be perturbed. BAM-ICL mainly consists of two stages: In the offline stage, where we assume the adversary has access to data drawn from the same distribution as the target task, we develop a global gradientbased attack to learn optimal budget allocations across ICEs. In the online stage, where ICEs arrive sequentially, perturbations are generated progressively according to the learned budget profile. We evaluate BAM-ICL on diverse LLMs and datasets, the experimental results demonstrate that it achieves superior attack success rates and stealthiness and the adversarial ICEs are highly transferable to other models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
相关 Paper
- Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context LearningShuai Zhao, Meihuizi Jia, Anh Tuan Luu, Fengjun Pan 等EMNLP 2024 · 被引用 28 次
- ICL-Evader: Zero-Query Black-Box Evasion Attacks on In-Context Learning and Their DefensesNingyuan He, Ronghong Huang, Qianqian Tang, Hongyu Wang 等WWW 2026
- Jailbreaking? One Step Is Enough!Weixiong Zheng, Peijian Zeng, Yiwei Li, Hongyan Wu 等ACL 2025 · 被引用 5 次
- Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt InjectionMeng Chen, Kun Wang, Li Lu, Jiaheng Zhang 等S&P 2026 · 被引用 3 次
- Phi: Preference Hijacking in Multi-modal Large Language Models at Inference TimeYifan Lan, Yuanpu Cao, Weitong Zhang, Lu Lin 等EMNLP 2025
