Robustness Auditing for Linear Regression: To Singularity and Beyond
Ittai Rubinstein, Samuel B. Hopkins
Abstract
It has recently been discovered that the conclusions of many highly influential econometrics studies can be overturned by removing a very small fraction of their samples (often less than 0.5%). These conclusions are typically based on the results of one or more Ordinary Least Squares (OLS) regressions, raising the question: given a dataset, can we certify the robustness of an OLS fit on this dataset to the removal of a given number of samples? Brute-force techniques quickly break down even on small datasets. Existing approaches which go beyond brute force either can only find candidate small subsets to remove (but cannot certify their non-existence) Broderick et al. ( 2020 ); Kuschnig et al. (2021), are computationally intractable beyond low dimensional settings Moitra & Rohatgi (2022), or require very strong assumptions on the data distribution and too many samples to give reasonable bounds in practice Bakshi & Prasad (2021); Freund & Hopkins ( 2023 ). We present an efficient algorithm for certifying the robustness of linear regressions to removals of samples. We implement our algorithm and run it on several landmark econometrics datasets with hundreds of dimensions and tens of thousands of samples, giving the first non-trivial certificates of robustness to sample removal for datasets of dimension 4 or greater. We prove that under distributional assumptions on a dataset, the bounds produced by our algorithm are tight up to a 1 + o(1) multiplicative factor. Published as a conference paper at ICLR 2025 Context and Applications for Robustness Auditing Problem 1 was introduced in this form by Broderick, Giordano, and Meager Broderick et al. (2020) , who use a heuristic algorithm, AMIP, to identify very small subsets of landmark datasets from econometrics which can be removed to overturn important conclusions of the respective studies Finkelstein et al. (2012); Angelucci & De Giorgi (2009) ; often this can be achieved by removing less than 0.5% of a dataset. Researchers have subsequently used AMIP to audit a wide range of recent studies in economics Martinez (2022); Di
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 546ae1cc-6088-414e-9f01-8b9c4e4ac3f7Cited by top-tier papers5
- What Data Enables Optimal Decisions? An Exact Characterization for Linear OptimizationOmar Bennouna, Amine Bennouna, Saurabh Amin, Asuman OzdaglarNeurIPS 2025 · 11 citations
- Rescaled Influence Functions: Accurate Data Attribution in High DimensionIttai Rubinstein, Samuel B. HopkinsNeurIPS 2025 · 3 citations
- Testing Most Influential SetsLucas Darius Konrad, Nikolas KuschnigICLR 2026 · 3 citations
- Finding Most Influential SetsLucas D. Konrad, Nikolas KuschnigICML 2026
- On the Accuracy of Newton Step and Influence Function Data AttributionsIttai Rubinstein, Samuel HopkinsICML 2026
Builds on2
Related papers
- Stress-Testing Causal Claims via Cardinality RepairsYarden Gabbay, Haoquan Guan, Shaull Almagor, El Kindi Rezig et al.SIGMOD 2026 · 1 citation
- Robust Testing in High-Dimensional Sparse ModelsAnand Jerry George, Clément L. CanonneNeurIPS 2022 · 4 citations
- List Decodable Mean Estimation in Nearly Linear TimeYeshwanth Cherapanamjeri, Sidhanth Mohanty, Morris YauFOCS 2020 · 13 citations
- Active Linear Regression for ℓp Norms and BeyondCameron Musco, Christopher Musco, David P. Woodruff, Taisuke YasudaFOCS 2022 · 4 citations
- SOL: Sampling-based Optimal Linear bounding of arbitrary scalar functionsYuriy Biktairov, Jyotirmoy DeshmukhNeurIPS 2023 · 3 citations
