Your Weak LLM is Secretly a Strong Teacher for Alignment
Leitian Tao, Yixuan Li
Abstract
The burgeoning capabilities of large language models (LLMs) have underscored the need for alignment to ensure these models act in accordance with human values and intentions. Existing alignment frameworks present constraints either in the form of expensive human effort or high computational costs. This paper explores a promising middle ground, where we employ a weak LLM that is significantly less resource-intensive than top-tier models, yet offers more automation than purely human feedback. We present a systematic study to evaluate and understand weak LLM's ability to generate feedback for alignment. Our empirical findings demonstrate that weak LLMs can provide feedback that rivals or even exceeds that of fully human-annotated data. Our study indicates a minimized impact of model size on feedback efficacy, shedding light on a scalable and sustainable alignment strategy. To deepen our understanding of alignment under weak LLM feedback, we conduct a series of qualitative and quantitative analyses, offering novel insights into the quality discrepancies between human feedback vs. weak LLM feedback.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c748ca80-ee1e-4bd0-a6ef-ebae49a89614Cited by top-tier papers12
- Towards Acyclic Preference Evaluation of Language Models via Multiple EvaluatorsZhengyu Hu, Jieyu Zhang, Zhihan Xiong, Alexander Ratner et al.AAAI 2026 · 14 citations
- On the Mechanisms of Weak-to-Strong Generalization: A Theoretical PerspectiveBehrad Moniri, Hamed HassaniNeurIPS 2025 · 8 citations
- Sparta Alignment: Collectively Aligning Multiple Language Models through CombatYuru Jiang, Wenxuan Ding, Shangbin Feng, Greg Durrett et al.NeurIPS 2025 · 8 citations
- Can DPO Learn Diverse Human Values? A Theoretical Scaling LawShawn Im, Sharon LiNeurIPS 2025 · 8 citations
- Contrastive Weak-to-Strong GeneralizationHoucheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang et al.ICML 2026 · 2 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
Related papers
- When Weak LLMs Speak with Confidence, Preference Alignment Gets StrongerAmirabbas Afzali, Myeongho Jeon, Maria BrbicICLR 2026
- Weak to Strong Generalization for Large Language Models with Multi-capabilitiesYucheng Zhou, Jianbing Shen, Yu ChengICLR 2025
- Aligning Large Language Models through Synthetic FeedbackSungdong Kim, Sanghwan Bae, Jamin Shin, Soyoung Kang et al.EMNLP 2023 · 13 citations
- A transfer learning framework for weak to strong generalizationSeamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Yaacov Ritov et al.ICLR 2025
- Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language ModelsHritik Bansal, John Dang, Aditya GroverICLR 2024 · 28 citations
