The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
Peter Hase, Mohit Bansal, Peter Clark, Sarah Wiegreffe
Abstract
How can we train models to perform well on hard test data when hard training data is by definition difficult to label correctly? This question has been termed the scalable oversight problem and has drawn increasing attention as language models have continually improved. In this paper, we present the surprising conclusion that current pretrained language models often generalize relatively well from easy to hard data, even performing as well as oracle models finetuned on hard data. We demonstrate this kind of easy-to-hard generalization using simple finetuning methods like in-context learning, linear classifier heads, and QLoRA for seven different measures of datapoint hardness, including six empirically diverse human hardness measures (like grade level) and one model-based measure (loss-based). Furthermore, we show that even if one cares most about model performance on hard data, it can be better to collect easy data rather than hard data for finetuning, since hard data is generally noisier and costlier to collect. Our experiments use open models up to 70b in size and four publicly available question-answering datasets with questions ranging in difficulty from 3rd grade science questions to college level STEM questions and general-knowledge trivia. We conclude that easy-to-hard generalization in LMs is surprisingly strong for the tasks studied. 1 Test Input LM Generated Answer Q: John hires a driving service to get him to work each day. His work is 30 miles away and he has to go there and back each day. He goes to work 5 days a week for 50 weeks a year. He gets charged 150 bonus per month How much does he pay a year for driving? A: John goes to work 5 days a week for 50 weeks a year. John goes to work 5 x 50 = <<550=250>>250 times a year. John pays 2 x 30 x 2 = <<230*2=120>>120 for each trip.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6d6d74d-23bb-404b-84e3-c32153831d74Cited by top-tier papers17
- Easy-to-Hard Generalization: Scalable Alignment Beyond Human SupervisionZhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu et al.NeurIPS 2024 · 125 citations
- Can Language Models Learn to Skip Steps?Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang et al.NeurIPS 2024 · 92 citations
- AI Debate Aids Assessment of Controversial ClaimsSalman Rahman, Sheriff Issaka, Ashima Suvarna, Genglin Liu et al.NeurIPS 2025 · 10 citations
- Beyond Oracle: Verifier-Supervision for Instruction Hierarchy in Reasoning and Instruction-Tuned LLMsSian-Yao Huang, Li-Hsien Chang, Che-Yu Lin, Cheng-Lin YangNeurIPS 2025 · 4 citations
- Can Large Language Models Generalize Procedures Across Representations?Fangru Lin, Valentin Hofmann, Xingchen Wan, Weixing Wang et al.ICML 2026 · 2 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
Related papers
- Physics of Language Models: Part 3.2, Knowledge ManipulationZeyuan Allen-Zhu, Yuanzhi LiICLR 2025 · 2 citations
- What Do Learning Dynamics Reveal About Generalization in LLM Mathematical Reasoning?Katie Kang, Amrith Setlur, Dibya Ghosh, Jacob Steinhardt et al.ICML 2025
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization ChallengesNayoung Lee, Ziyang Cai, Avi Schwarzschild, Kangwook Lee et al.ICML 2025
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
- Rapid Word Learning Through Meta In-Context LearningWentao Wang, Guangyuan Jiang, Tal Linzen, Brenden M. LakeEMNLP 2025
