Learning Preference Model for LLMs via Automatic Preference Data Generation
Shijia Huang, Jianqiao Zhao, Yanyang Li, Liwei Wang
摘要
Despite the advanced capacities of the state-of-the-art large language models (LLMs), they suffer from issues of hallucination, stereotype, etc. Preference models play an important role in LLM alignment, yet training preference models predominantly rely on human-annotated data. This reliance limits their versatility and scalability. In this paper, we propose learning the preference model for LLMs via automatic preference data generation (AutoPM). Our approach involves both In-Breadth Data Generation, which elicits pairwise preference data from LLMs following the helpful-honest-harmless (HHH) criteria, and In-Depth Data Generation, which enriches the dataset with responses spanning a wide quality range. With HHH-guided preference data, our approach simultaneously enables the LLMs to learn human preferences and align with human values. Quantitative assessments on five benchmark datasets demonstrate the reliability and potential of AutoPM, pointing out a more general and scalable way to improve LLM performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Semi-Supervised Preference Optimization with Limited FeedbackSeonggyun Lee, Sungjun Lim, Seojin Park, Soeun Cheon 等ICLR 2026 · 被引用 4 次
- Focus On This, Not That! Steering LLMs with Adaptive Feature SpecificationTom A. Lamb, Adam Davies, Alasdair Paren, Philip Torr 等ICML 2025
- SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation SteeringUtsav Maskey, Sumit Yadav, Mark Dras, Usman NaseemACL 2026
- Improving Reward Model Generalization from Adversarial Process Enhanced PreferencesZhilong Zhang, Tian Xu, Xinghao Du, Xingchen Cao 等ICML 2025
- Measuring Diversity in Synthetic DatasetsYuchang Zhu, Huizhe Zhang, Bingzhe Wu, Jintang Li 等ICML 2025
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang 等NeurIPS 2023 · 被引用 463 次
相关 Paper
- ActiveDPO: Active Direct Preference Optimization for Sample-Efficient AlignmentXiaoqiang Lin, Arun Verma, Zhongxiang Dai, Daniela Rus 等ICLR 2026 · 被引用 12 次
- Larger or Smaller Reward Margins to Select Preferences for LLM Alignment?Kexin Huang, Junkang Wu, Ziqian Chen, Xue Wang 等ICML 2025
- Data Selection for LLM Alignment Using Fine-Grained PreferencesJia Zhang, Yao Liu, Chen-Xi Zhang, Yi Liu 等ICLR 2026 · 被引用 1 次
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong 等ACL 2024 · 被引用 4 次
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen 等NeurIPS 2025 · 被引用 13 次
