Learning from Biased Data: A Semi-Parametric Approach
Patrice Bertail, Stéphan Clémençon, Yannick Guyonvarch, Nathan Noiry
摘要
We consider risk minimization problems where the (source) distribution P S of the training observations Z 1 , . . . , Z n differs from the (target) distribution P T involved in the risk that one seeks to minimize. Under the natural assumption that P S dominates P T , i.e. P T < < P S , we develop a semiparametric framework in the situation where we do not observe any sample from P T , but rather have access to some auxiliary information at the target population scale. More precisely, assuming that the Radon-Nikodym derivative dP T /dP S (z) belongs to a parametric class g(z, α), α ∈ A and that some (generalized) moments of P T are available to the learner, we propose a two-step learning procedure to perform the risk minimization task. We first select α so as to match the moment constraints as closely as possible and then reweight each (biased) training observation Z i by g(Z i , α) in the final Empirical Risk Minimization (ERM) algorithm. We establish a O P (1/ √ n) generalization bound proving that, remarkably, the solution to the weighted ERM problem thus constructed achieves a learning rate of the same order as that attained in absence of any sampling bias. Beyond these theoretical guarantees, numerical results providing strong empirical evidence of the relevance of the approach promoted in this article are displayed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Off-Policy Evaluation with Policy-Dependent Optimization ResponseWenshuo Guo, Michael I. Jordan, Angela ZhouNeurIPS 2022 · 被引用 5 次
- Best Arm Identification with Biased ContextsJames Cheshire, Stéphan ClémençonAAAI 2026
它引用的顶会 Paper1
相关 Paper
- Optimal Excess Risk Bounds for Empirical Risk Minimization on p-Norm Linear RegressionAyoub El Hanchi, Murat A. ErdogduNeurIPS 2023 · 被引用 2 次
- Double-Weighting for Covariate Shift AdaptationJosé Ignacio Segovia-Martín, Santiago Mazuelas, Anqi LiuICML 2023 · 被引用 9 次
- Scaling laws for learning with real and surrogate dataAyush Jain, Andrea Montanari, Eren SasogluNeurIPS 2024 · 被引用 30 次
- Risk Minimization from Adaptively Collected Data: Guarantees for Supervised and Policy LearningAurélien Bibaut, Nathan Kallus, Maria Dimakopoulou, Antoine Chambaz 等NeurIPS 2021 · 被引用 18 次
- A Distribution-dependent Analysis of Meta LearningMikhail Konobeev, Ilja Kuzborskij, Csaba SzepesváriICML 2021 · 被引用 6 次
