Robustness Auditing for Linear Regression: To Singularity and Beyond
Ittai Rubinstein, Samuel B. Hopkins
摘要
It has recently been discovered that the conclusions of many highly influential econometrics studies can be overturned by removing a very small fraction of their samples (often less than 0.5%). These conclusions are typically based on the results of one or more Ordinary Least Squares (OLS) regressions, raising the question: given a dataset, can we certify the robustness of an OLS fit on this dataset to the removal of a given number of samples? Brute-force techniques quickly break down even on small datasets. Existing approaches which go beyond brute force either can only find candidate small subsets to remove (but cannot certify their non-existence) Broderick et al. ( 2020 ); Kuschnig et al. (2021), are computationally intractable beyond low dimensional settings Moitra & Rohatgi (2022), or require very strong assumptions on the data distribution and too many samples to give reasonable bounds in practice Bakshi & Prasad (2021); Freund & Hopkins ( 2023 ). We present an efficient algorithm for certifying the robustness of linear regressions to removals of samples. We implement our algorithm and run it on several landmark econometrics datasets with hundreds of dimensions and tens of thousands of samples, giving the first non-trivial certificates of robustness to sample removal for datasets of dimension 4 or greater. We prove that under distributional assumptions on a dataset, the bounds produced by our algorithm are tight up to a 1 + o(1) multiplicative factor. Published as a conference paper at ICLR 2025 Context and Applications for Robustness Auditing Problem 1 was introduced in this form by Broderick, Giordano, and Meager Broderick et al. (2020) , who use a heuristic algorithm, AMIP, to identify very small subsets of landmark datasets from econometrics which can be removed to overturn important conclusions of the respective studies Finkelstein et al. (2012); Angelucci & De Giorgi (2009) ; often this can be achieved by removing less than 0.5% of a dataset. Researchers have subsequently used AMIP to audit a wide range of recent studies in economics Martinez (2022); Di
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- What Data Enables Optimal Decisions? An Exact Characterization for Linear OptimizationOmar Bennouna, Amine Bennouna, Saurabh Amin, Asuman OzdaglarNeurIPS 2025 · 被引用 11 次
- Rescaled Influence Functions: Accurate Data Attribution in High DimensionIttai Rubinstein, Samuel B. HopkinsNeurIPS 2025 · 被引用 3 次
- Testing Most Influential SetsLucas Darius Konrad, Nikolas KuschnigICLR 2026 · 被引用 3 次
- Finding Most Influential SetsLucas D. Konrad, Nikolas KuschnigICML 2026
- On the Accuracy of Newton Step and Influence Function Data AttributionsIttai Rubinstein, Samuel HopkinsICML 2026
它引用的顶会 Paper2
相关 Paper
- Stress-Testing Causal Claims via Cardinality RepairsYarden Gabbay, Haoquan Guan, Shaull Almagor, El Kindi Rezig 等SIGMOD 2026 · 被引用 1 次
- Robust Testing in High-Dimensional Sparse ModelsAnand Jerry George, Clément L. CanonneNeurIPS 2022 · 被引用 4 次
- List Decodable Mean Estimation in Nearly Linear TimeYeshwanth Cherapanamjeri, Sidhanth Mohanty, Morris YauFOCS 2020 · 被引用 13 次
- Active Linear Regression for ℓp Norms and BeyondCameron Musco, Christopher Musco, David P. Woodruff, Taisuke YasudaFOCS 2022 · 被引用 4 次
- SOL: Sampling-based Optimal Linear bounding of arbitrary scalar functionsYuriy Biktairov, Jyotirmoy DeshmukhNeurIPS 2023 · 被引用 3 次
