PICACO: Pluralistic In-Context Value Alignment via Total Correlation Optimization
Han Jiang, Dongyao Zhu, Xiaoyuan Yi, Ziang Xiao, Zhihua Wei, Xing Xie
摘要
In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and accommodate diverse preferences without costly post-training, known as In-Context Alignment (ICA). However, LLMs' comprehension of input prompts remains agnostic, limiting ICA's ability to address value tensions-human values are inherently pluralistic, often imposing conflicting demands, e.g., stimulation vs. tradition. Current ICA methods therefore face the Instruction Bottleneck challenge, where LLMs struggle to reconcile multiple intended values within a single prompt, leading to incomplete or biased alignment. To address this, we propose PICACO, a novel pluralistic ICA method. Without fine-tuning, PI-CACO optimizes a meta-instruction that incorporates multiple values to better elicit LLMs' understanding of them and improve alignment. This is achieved by maximizing the total correlation between specified values and LLM responses, which theoretically reinforces value conformity and reduces distractive noise, resulting in more effective instructions. Extensive experiments on five value sets show that PICACO works well with both black-box and open-source LLMs, outperforms several recent strong baselines, and achieves a better balance across up to 8 distinct values. Being free means having the autonomy to make choices and decisions without undue restrictions, allowing individuals to pursue their passions, express themselves authentically, and live life according to their own values and beliefs. Please provide responses that better reflect the values of Stimulation, Hedonism, Universalism, and Self-direction. What does being free mean to you? I'm sorry, I can't assist with that request.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper42
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri 等NeurIPS 2023 · 被引用 516 次
- RRHF: Rank Responses to Align Language Models with Human FeedbackHongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang 等NeurIPS 2023 · 被引用 515 次
相关 Paper
- Multi-Value Alignment for LLMs via Value Decorrelation and ExtrapolationHefei Xu, Le Wu, Chen Cheng, Hao LiuAAAI 2026 · 被引用 4 次
- Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language ModelsJongwook Han, Jongwon Lim, Injin Kong, Yohan JoICML 2026
- Do LLMs have Consistent Values?Naama Rozen, Liat Bezalel, Gal Elidan, Amir Globerson 等ICLR 2025
- Aligning to Thousands of Preferences via System Message GeneralizationSeongyun Lee, Sue Hyun Park, Seungone Kim, Minjoon SeoNeurIPS 2024 · 被引用 102 次
- PICLe: Eliciting Diverse Behaviors from Large Language Models with Persona In-Context LearningHyeong Kyu Choi, Yixuan LiICML 2024 · 被引用 31 次
