Datamodels: Understanding Predictions with Data and Data with Predictions
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, Aleksander Madry
摘要
We present a conceptual framework, datamodeling, for analyzing the behavior of a model class in terms of the training data. For any fixed "target" example x, training set S, and learning algorithm, a datamodel is a parameterized function 2 S → R that for any subset of S ⊂ Susing only information about which examples of S are contained in S -predicts the outcome of training a model on S and evaluating on x. Despite the complexity of the underlying process that is being approximated (e.g. end-to-end training and evaluation of deep neural networks), we show that even simple linear datamodels successfully predict model outputs. We then demonstrate that datamodels give rise to a variety of applications, such as: accurately predicting the effect of dataset counterfactuals; identifying brittle predictions; finding semantically similar examples; quantifying train-test leakage; and embedding data into a well-behaved and feature-rich representation space.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James ZouICML 2023 · 被引用 54 次
- TSDS: Data Selection for Task-Specific Model FinetuningZifan Liu, Amin Karbasi, Theodoros RekatsinasNeurIPS 2024 · 被引用 35 次
- Hubble: a Model Suite to Advance the Study of LLM MemorizationJohnny Wei, Ameya Godbole, Mohammad Aflah Khan, Ryan Yixiang Wang 等ICLR 2026 · 被引用 22 次
- ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMsLandon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas 等NeurIPS 2025 · 被引用 19 次
- Robust Weight Signatures: Gaining Robustness as Easy as Patching Weights?Ruisi Cai, Zhenyu Zhang, Zhangyang WangICML 2023 · 被引用 16 次
它引用的顶会 Paper15
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
相关 Paper
- Machine Learning Models that Remember Too MuchCongzheng Song, Thomas Ristenpart, Vitaly ShmatikovCCS 2017 · 被引用 582 次
- Estimating informativeness of samples with Smooth Unique InformationHrayr Harutyunyan, Alessandro Achille, Giovanni Paolini, Orchid Majumder 等ICLR 2021 · 被引用 26 次
- Distilled Datamodel with Reverse Gradient MatchingJingwen Ye, Ruonan Yu, Songhua Liu, Xinchao WangCVPR 2024 · 被引用 1 次
- From Black-box to Causal-box: Towards Building More Interpretable ModelsInwoo Hwang, Yushu Pan, Elias BareinboimNeurIPS 2025 · 被引用 3 次
- Counterfactual Memorization in Neural Language ModelsChiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski 等NeurIPS 2023 · 被引用 184 次
