How to Mitigate Overfitting in Weak-to-strong Generalization?
Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng, Qipeng Guo, Xipeng Qiu
Abstract
Aligning powerful AI models on tasks that surpass human evaluation capabilities is the central problem of superalignment. To address this problem, weak-to-strong generalization aims to elicit the capabilities of strong models through weak supervisors and ensure that the behavior of strong models aligns with the intentions of weak supervisors without unsafe behaviors such as deception. Although weak-to-strong generalization exhibiting certain generalization capabilities, strong models exhibit significant overfitting in weak-to-strong generalization: Due to the strong fit ability of strong models, erroneous labels from weak supervisors may lead to overfitting in strong models. In addition, simply filtering out incorrect labels may lead to a degeneration in question quality, resulting in a weak generalization ability of strong models on hard questions. To mitigate overfitting in weak-to-strong generalization, we propose a two-stage framework that simultaneously improves the quality of supervision signals and the quality of input questions. Experimental results in three series of large language models and two mathematical benchmarks demonstrate that our framework significantly improves PGR compared to naive weak-to-strong generalization, even achieving up to 100% PGR on some models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5607a10-9574-4bf6-ab70-867fc7e9af32Builds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Theoretical Analysis of Weak-to-Strong GeneralizationHunter Lang, David A. Sontag, Aravindan VijayaraghavanNeurIPS 2024 · 59 citations
- The Unreasonable Effectiveness of Easy Training Data for Hard TasksPeter Hase, Mohit Bansal, Peter Clark, Sarah WiegreffeACL 2024 · 3 citations
- MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference OptimizationYougang Lyu, Lingyong Yan, Zihan Wang, Dawei Yin et al.ICLR 2025
Related papers
- Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong GeneralizationWenkai Yang, Shiqi Shen, Guangyao Shen, Wei Yao et al.ICLR 2025
- Weak to Strong Generalization for Large Language Models with Multi-capabilitiesYucheng Zhou, Jianbing Shen, Yu ChengICLR 2025
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- A transfer learning framework for weak to strong generalizationSeamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Yaacov Ritov et al.ICLR 2025
- Quantifying the Gain in Weak-to-Strong GeneralizationMoses Charikar, Chirag Pabbaraju, Kirankumar ShiragurNeurIPS 2024 · 42 citations
