Small-to-Large Generalization: Training Data Influences Models Consistently Across Scale
Alaa Khaddaj, Logan Engstrom, Aleksander Madry
摘要
Choice of training data distribution greatly influences model behavior. Yet, in large-scale settings, precisely characterizing how changes in training data affects predictions is often difficult due to model training costs. Current practice is to instead extrapolate from scaled down, inexpensive-to-train proxy models. However, changes in data do not influence smaller and larger models identically. Therefore, understanding how choice of data affects large-scale models raises the question: how does training data distribution influence model behavior across compute scale? We find that small-and large-scale language model predictions (generally) do highly correlate across choice of training data. Equipped with these findings, we characterize how proxy scale affects effectiveness in two downstream proxy model applications: data attribution and dataset selection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- IF-Guide: Influence Function-Guided Detoxification of LLMsZachary Coalson, Juhan Bae, Nicholas Carlini, Sanghyun HongNeurIPS 2025 · 被引用 9 次
- Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model PracticeJiachen T. Wang, Tong Wu, Kaifeng Lyu, James Zou 等ICLR 2026 · 被引用 3 次
- A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn’t)Nihal Nayak, Paula Rodriguez-Diaz, Neha Hulkund, Sara Beery 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
相关 Paper
- AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context AttributionFengyuan Liu, Nikhil Kandpal, Colin RaffelICLR 2025
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman 等ICLR 2026 · 被引用 8 次
- DataDecide: How to Predict Best Pretraining Data with Small ExperimentsIan Magnusson, Nguyen Tai, Ben Bogin, David Heineman 等ICML 2025
- A Hitchhiker's Guide to Scaling Law EstimationLeshem Choshen, Yang Zhang, Jacob AndreasICML 2025
- LLMs on the Line: Data Determines Loss-to-Loss Scaling LawsPrasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge 等ICML 2025
