Harnessing Non-Adversarial Robustness in Large Language Models
Qinghua Zhou, Ellina Aleshina, Andrey Lovyagin, Oleg Somov, Mikhail Seleznyov, Alexander Panchenko, Ivan Oseledets, Elena Tutubalina, Ivan Tyukin
摘要
The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by semantically similar but textually different prompts. Recent works have shown that these kinds of prompt variations can significantly impact the performance of LLMs on tasks. The central question is: can LLMs' robustness to semantically-neutral prompt alterations be acquired without expensive retraining of the entire model? We address this question both theoretically and through experiments. Our theoretical analysis reveals a crucial factor impacting model robustness -- a systematic expected shift or perturbation-induced bias in neural network module outputs. Motivated by this analysis, we show that robustness can be achieved via a simple fine-tuning process: debiasing for robustness. We identify conditions when debiasing helps and when it does not, and demonstrate, through both theory and extensive experiments, that debiasing for robustness may indeed be a quick and efficient tool to enhance robustness and provide certification against random prompt perturbations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 被引用 208 次
- Efficient Adversarial Training in LLMs with Continuous AttacksSophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel 等NeurIPS 2024 · 被引用 151 次
- Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt EngineeringHan Zhou, Xingchen Wan, Lev Proleev, Diana Mincu 等ICLR 2024 · 被引用 90 次
相关 Paper
- Evaluating Concurrent Robustness of Language Models Across Diverse Challenge SetsVatsal Gupta, Pranshu Pandya, Tushar Kataria, Vivek Gupta 等EMNLP 2024 · 被引用 1 次
- Evaluating the Zero-shot Robustness of Instruction-tuned Language ModelsJiuding Sun, Chantal Shaib, Byron C. WallaceICLR 2024 · 被引用 75 次
- PAFT: Prompt-Agnostic Fine-TuningChenxing Wei, Mingwen Ou, Ying He, Yao Shu 等EMNLP 2025
- Prompt Perturbation in Retrieval-Augmented Generation based Large Language ModelsZhibo Hu, Chen Wang, Yanfeng Shu, Hye-Young Paik 等KDD 2024 · 被引用 15 次
- Evaluating and Explaining Prompt Sensitivity of LLMs Using InteractionsRuiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei 等ICML 2026 · 被引用 1 次
