InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
Pengyu Wang, Dong Zhang, Linyang Li, Chenkun Tan, Xinghao Wang, Mozhi Zhang, Ke Ren, Botian Jiang, Xipeng Qiu
摘要
As large language models (LLMs) rapidly evolve, they are increasingly being customized through fine-tuning to suit the specific needs of various applications. A critical aspect of this advancement is the alignment process, which ensures that these models perform tasks in ways that align with human values and expectations. Current alignment methods, such as direct preference optimization (DPO) and reinforcement learning from human feedback (RLHF), focus primarily on alignment during training phase. However, these methods often involve complex and resource-intensive training processes, posing significant challenge for their implementation. Therefore, we propose InferAligner, a simple yet effective method for harmlessness alignment during inference phase. InferAligner decouples harmlessness from helpfulness. During the training phase, it focuses solely on enhancing the target model's capabilities on downstream tasks. In the inference phase, it utilizes safety steering vectors extracted from the aligned model to guide the target model towards harmlessness alignment. Experimental results show that our method can be very effectively applied to domain-specific models in finance, medicine, and mathematics, as well as to multimodal large language models (MLLMs) such as LLaVA. It significantly diminishes the attack success rate (ASR) of both harmful instructions and jailbreak instructions, while maintaining almost unchanged performance in downstream tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Fast Best-of-N Decoding via Speculative RejectionHanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang 等NeurIPS 2024 · 被引用 144 次
- SpeechAlign: Aligning Speech Generation to Human PreferencesDong Zhang, Zhaowei Li, Shimin Li, Xin Zhang 等NeurIPS 2024 · 被引用 74 次
- GuardReasoner-VL: Safeguarding VLMs via Reinforced ReasoningYue Liu, Shengfang Zhai, Mingzhe Du, Yulin Chen 等NeurIPS 2025 · 被引用 40 次
- The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and DefenseYangyang Guo, Fangkai Jiao, Liqiang Nie, Mohan KankanhalliNeurIPS 2025 · 被引用 24 次
- Understanding and Rectifying Safety Perception Distortion in VLMsXiaohan Zou, Jian Kang, George Kesidis, Lu LinNeurIPS 2025 · 被引用 20 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement LearningYi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng 等ICLR 2026 · 被引用 15 次
- Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak AttacksYingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng 等NDSS 2026 · 被引用 5 次
- Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu 等ACL 2024
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
- SDA: Steering-Driven Distribution Alignment for Open LLMs Without Fine-TuningWei Xia, Zhi-Hong DengAAAI 2026
