Self-Guided Alignment: Adaptive Preference Sensing for Multi-Objective Generation
Ning Wang, Zhanyang Liu, Taotao Zhou, Xinrui Zhang, Zongru Shao, Haojie Zhou
Abstract
Aligning Large Language Models (LLMs) with diverse and potentially conflicting human values necessitates navigating complex multi-objective landscapes. However, existing prompt-conditioned approaches face a critical training-inference discrepancy: they rely on ground-truth scores during training while requiring manual user-specification at inference. We introduce prediction of implicit preferences to bridge this gap while reducing user burden. To this end, we propose Self-Guided Alignment (SGA), a framework that transforms passive reward dependency into an intrinsic adaptive sensing capability. It employs a dual-head architecture to unify preference internalization with conditional generation, enabling the model to learn a latent mapping between raw prompts and preference profiles. Through adaptive preference sensing, the model autonomously predicts the latent preference score to self-guide the generation, thereby eliminating the need for manual specification at inference. Extensive experiments across diverse model scales demonstrate that SGA often outperforms state-of-the-art baselines, achieving superior multi-objective trade-offs and improved preference alignment. Code is available at https://github.com/python-yyds/SGA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ce2a3c4-c88c-4463-8731-273bdc4ea38cBuilds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsRui Yang, Ruomeng Ding, Yong Lin, Huan Zhang et al.NeurIPS 2024 · 157 citations
Related papers
- MetaAligner: Towards Generalizable Multi-Objective Alignment of Language ModelsKailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang et al.NeurIPS 2024 · 49 citations
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong et al.ACL 2024 · 4 citations
- Larger or Smaller Reward Margins to Select Preferences for LLM Alignment?Kexin Huang, Junkang Wu, Ziqian Chen, Xue Wang et al.ICML 2025
- Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language ModelsSomanshu Singla, Zhen Wang, Tianyang Liu, Abdullah Ashfaq et al.EMNLP 2024 · 1 citation
- IPO: Your Language Model is Secretly a Preference ClassifierShivank Garg, Ayush Singh, Shweta Singh, Paras ChopraACL 2025
