Scale-invariant Optimal Sampling for Rare-events Data and Sparse Models
Jing Wang, HaiYing Wang, Hao Zhang
摘要
Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect estimation efficiency, and optimal subsampling is essential to mitigate the information loss. However, existing optimal subsampling probabilities depends on data scales, and some scaling transformations may result in inefficient subsamples. This problem is more significant when there are inactive features, because their influence on the subsampling probabilities can be arbitrarily magnified by inappropriate scaling transformations. We tackle this challenge and introduce a scale-invariant optimal subsampling function in the context of sparse models, where inactive features are commonly assumed. Instead of focusing on estimating model parameters, we define an optimal subsampling function to minimize the prediction error, using adaptive lasso as an example to outline the estimation procedure and study its theoretical guarantee. We first introduce the adaptive lasso estimator for rare-events data and establish its oracle properties, thereby validating the use of subsampling. Then we derive a scale-invariant optimal subsampling function that minimizes the prediction error of the inverse probability weighted (IPW) adaptive lasso. Finally, we present an estimator based on the maximum sampled conditional likelihood (MSCL) to further improve the estimation efficiency. We conduct numerical experiments using both simulated and real-world data sets to demonstrate the performance of the proposed methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Optimal Sampling Gaps for Adaptive Submodular MaximizationShaojie Tang, Jing YuanAAAI 2022 · 被引用 10 次
- Robust Gradient-Based Markov SubsamplingTieliang Gong, Quanhan Xi, Chen XuAAAI 2020 · 被引用 3 次
- Towards a statistical theory of data selection under weak supervisionGermain Kolossov, Andrea Montanari, Pulkit TandonICLR 2024 · 被引用 27 次
- Less Is Better: Unweighted Data Subsampling via Influence FunctionZifeng Wang, Hong Zhu, Zhenhua Dong, Xiuqiang He 等AAAI 2020 · 被引用 61 次
- Scalable Counterfactual Risk Estimation for Rare Events in Longitudinal DataXiaohui Yin, Avijit Mitra, Ying Zhou, Kun Chen 等KDD 2026
