Lune

ACL2024顶会

The Unreasonable Effectiveness of Easy Training Data for Hard Tasks

Peter Hase, Mohit Bansal, Peter Clark, Sarah Wiegreffe

2024年份
3被引次数
17顶会引用

摘要

How can we train models to perform well on hard test data when hard training data is by definition difficult to label correctly? This question has been termed the scalable oversight problem and has drawn increasing attention as language models have continually improved. In this paper, we present the surprising conclusion that current pretrained language models often generalize relatively well from easy to hard data, even performing as well as oracle models finetuned on hard data. We demonstrate this kind of easy-to-hard generalization using simple finetuning methods like in-context learning, linear classifier heads, and QLoRA for seven different measures of datapoint hardness, including six empirically diverse human hardness measures (like grade level) and one model-based measure (loss-based). Furthermore, we show that even if one cares most about model performance on hard data, it can be better to collect easy data rather than hard data for finetuning, since hard data is generally noisier and costlier to collect. Our experiments use open models up to 70b in size and four publicly available question-answering datasets with questions ranging in difficulty from 3rd grade science questions to college level STEM questions and general-knowledge trivia. We conclude that easy-to-hard generalization in LMs is surprisingly strong for the tasks studied. 1 Test Input LM Generated Answer Q: John hires a driving service to get him to work each day. His work is 30 miles away and he has to go there and back each day. He goes to work 5 days a week for 50 weeks a year. He gets charged 2permiledrivenandhealsogiveshisdrivera2 per mile driven and he also gives his driver a 150 bonus per month How much does he pay a year for driving? A: John goes to work 5 days a week for 50 weeks a year. John goes to work 5 x 50 = <<550=250>>250 times a year. John pays 2permiledriven.Johnpays2 per mile driven. John pays 2 x 30 x 2 = <<230*2=120>>120 for each trip.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper17

问问它们各自怎么用它

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖