Aligning to Thousands of Preferences via System Message Generalization
Seongyun Lee, Sue Hyun Park, Seungone Kim, Minjoon Seo
摘要
Although humans inherently have diverse values, current large language model (LLM) alignment methods often assume that aligning LLMs with the general public's preferences is optimal. A major challenge in adopting a more individualized approach to LLM alignment is its lack of scalability, as it involves repeatedly acquiring preference data and training new reward models and LLMs for each individual's preferences. To address these challenges, we propose a new paradigm where users specify what they value most within the system message, steering the LLM's generation behavior to better align with the user's intentions. However, a naive application of such an approach is non-trivial since LLMs are typically trained on a uniform system message (e.g.,"You are a helpful assistant") which limits their ability to generalize to diverse, unseen system messages. To improve this generalization, we create the Multifaceted Collection, a preference dataset with 192k combinations of values beyond generic helpfulness and harmlessness, spanning 65k user instructions. Using this dataset, we train a 7B LLM called Janus and test it on 921 prompts from 5 benchmarks (AlpacaEval 2.0, FLASK, Koala, MT-Bench, and Self-Instruct) by adding various unseen system messages that reflect user preferences. Janus achieves tie+win rate of 75.2%, 72.4%, and 66.4% against Mistral 7B Instruct v0.2, GPT-3.5 Turbo, and GPT-4, respectively. Unexpectedly, on three benchmarks focused on response helpfulness (AlpacaEval 2.0, MT-Bench, Arena Hard Auto v0.1), Janus also outperforms LLaMA 3 8B Instruct by a +4.0%, +0.1%, +3.0% margin, underscoring that training with a vast array of system messages could also enhance alignment to the general public's preference as well. Our code, dataset, benchmark, and models are available at https://github.com/kaistAI/Janus.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith 等ICLR 2026 · 被引用 41 次
- Improving Context-Aware Preference Modeling for Language ModelsSilviu Pitis, Ziang Xiao, Nicolas Le Roux, Alessandro SordoniNeurIPS 2024 · 被引用 30 次
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing 等ICLR 2026 · 被引用 29 次
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference DataRajiv Movva, Smitha Milli, Sewon Min, Emma PiersonICLR 2026 · 被引用 27 次
- Direct Alignment with Heterogeneous PreferencesAli Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri 等NeurIPS 2025 · 被引用 26 次
它引用的顶会 Paper37
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- Improving Model Alignment Through Collective Intelligence of Open-Source ModelsJunlin Wang, Roy Xie, Shang Zhu, Jue Wang 等ICML 2025
- Self-Boosting Large Language Models with Synthetic Preference DataQingxiu Dong, Li Dong, Xingxing Zhang, Zhifang Sui 等ICLR 2025
- Dissecting Human and LLM PreferencesJunlong Li, Fan Zhou, Shichao Sun, Yikai Zhang 等ACL 2024 · 被引用 1 次
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 被引用 3 次
- From 1, 000, 000 Users to Every User: Scaling Up Personalized Preference for User-level AlignmentJia-Nan Li, Jian Guan, Songhao Wu, Wei Wu 等ACL 2026
