Measure the Predictive Heterogeneity
Jiashuo Liu, Jiayun Wu, Renjie Pi, Renzhe Xu, Xingxuan Zhang, Bo Li, Peng Cui
Abstract
As an intrinsic and fundamental property of big data, data heterogeneity exists in a variety of real-world applications, such as in agriculture, sociology, health care, etc. For machine learning algorithms, the ignorance of data heterogeneity will significantly hurt the generalization performance and the algorithmic fairness, since the prediction mechanisms among different sub-populations are likely to differ. In this work, we focus on the data heterogeneity that affects the prediction of machine learning models, and first formalize the Predictive Heterogeneity, which takes into account the model capacity and computational constraints. We prove that it can be reliably estimated from finite data with PAC bounds even in high dimensions. Additionally, we propose the Information Maximization (IM) algorithm, a bi-level optimization algorithm, to explore the predictive heterogeneity of data. Empirically, the explored predictive heterogeneity provides insights for sub-population divisions in agriculture, sociology, and object recognition, and leveraging such heterogeneity benefits the out-of-distribution generalization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3f3bc4b-f1fe-49a1-bf69-def69d1645eaCited by top-tier papers3
- Quantitatively Measuring and Contrastively Exploring Heterogeneity for Domain GeneralizationYunze Tong, Junkun Yuan, Min Zhang, Didi Zhu et al.KDD 2023 · 7 citations
- Heterogeneous Data Game: Characterizing the Model Competition Across Multiple Data SourcesRenzhe Xu, Kang Wang, Bo LiICML 2025
- Rethinking the Evaluation Protocol of Domain GeneralizationHan Yu, Xingxuan Zhang, Renzhe Xu, Jiashuo Liu et al.CVPR 2024
Builds on5
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 454 citations
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart et al.ICLR 2020 · 211 citations
- Heterogeneous Risk MinimizationJiashuo Liu, Zheyuan Hu, Peng Cui, Bo Li et al.ICML 2021 · 170 citations
- Model Agnostic Sample Reweighting for Out-of-Distribution LearningXiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang et al.ICML 2022 · 73 citations
- Comparing Distributions by Measuring Differences that Affect Decision MakingShengjia Zhao, Abhishek Sinha, Yutong He, Aidan Perreault et al.ICLR 2022 · 28 citations
Related papers
- On Harmonizing Implicit SubpopulationsFeng Hong, Jiangchao Yao, Yueming Lyu, Zhihan Zhou et al.ICLR 2024 · 8 citations
- Multigroup RobustnessLunjia Hu, Charlotte Peale, Judy Hanwen ShenICML 2024 · 2 citations
- Applied Online Algorithms with Heterogeneous PredictorsJessica Maghakian, Russell Lee, Mohammad Hajiesmaili, Jian Li et al.ICML 2023 · 7 citations
- Data Exchange Markets via Utility BalancingAditya Bhaskara, Sreenivas Gollapudi, Sungjin Im, Kostas Kollias et al.WWW 2024 · 6 citations
- Discover and Mitigate Multiple Biased Subgroups in Image ClassifiersZeliang Zhang, Mingqian Feng, Zhiheng Li, Chenliang XuCVPR 2024 · 5 citations
