Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active Learning
Hugo Cui, Yue Lu
Abstract
We study a class of iterated empirical risk minimization (ERM) procedures in which two successive ERMs are performed on the same dataset, and the predictions of the first estimator enter as an argument in the loss function of the second. This setting, which arises naturally in active learning and reweighting schemes, introduces intricate statistical dependencies across samples and fundamentally distinguishes the problem from classical single-stage ERM analyses. For linear models trained with a broad class of convex losses on Gaussian mixture data, we derive a sharp asymptotic characterization of the test error in the high-dimensional regime where the sample size and ambient dimension scale proportionally. Our results provide explicit, fully asymptotic predictions for the performance of the second-stage estimator despite the reuse of data and the presence of prediction-dependent losses. We apply this theory to revisit a well-studied pool-based active learning problem, removing oracle and sample-splitting assumptions made in prior work. We uncover a fundamental tradeoff in how the labeling budget should be allocated across stages, and demonstrate a double-descent behavior of the test error driven purely by data selection, rather than model size or sample count.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64f3be1b-d42b-44c1-988a-87eaaa1ec1adBuilds on12
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 163 citations
- Convergence of Uncertainty Sampling for Active LearningAnant Raj, Francis R. BachICML 2022 · 41 citations
- Towards a statistical theory of data selection under weak supervisionGermain Kolossov, Andrea Montanari, Pulkit TandonICLR 2024 · 27 citations
- How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and BudgetGuy Hacohen, Daphna WeinshallNeurIPS 2023 · 23 citations
Related papers
- Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensionsBruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco et al.NeurIPS 2021 · 70 citations
- Achieving Minimax Rates in Pool-Based Batch Active LearningClaudio Gentile, Zhilei Wang, Tong ZhangICML 2022 · 16 citations
- Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk MinimizationMohamed Chiheb Yaakoubi, Cosme Louart, Malik TIOMOKO, Zhenyu LiaoICML 2026
- Scaling laws for learning with real and surrogate dataAyush Jain, Andrea Montanari, Eren SasogluNeurIPS 2024 · 30 citations
- Efficient Active Learning for Gaussian Process Classification by Error ReductionGuang Zhao, Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander et al.NeurIPS 2021 · 28 citations
