Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular Data
Yunze Tong, Fengda Zhang, Zihao Tang, Kaifeng Gao, Kai Huang, Pengfei Lyu, Jun Xiao, Kun Kuang
摘要
Machine learning models often perform well on tabular data by optimizing average prediction accuracy. However, they may underperform on specific subsets due to inherent biases in the training data, such as associations with noncausal features like demographic information. These biases lead to critical robustness issues as models may inherit or amplify them, resulting in poor performance where such misleading correlations do not hold. Existing mitigation methods have significant limitations: some require prior group labels, which are often unavailable, while others focus solely on the conditional distribution P (Y |X), upweighting misclassified samples without effectively balancing the overall data distribution P (X). To address these shortcomings, we propose a latent score-based reweighting framework. It leverages score-based models to capture the joint data distribution P (X, Y ) without relying on additional prior information. By estimating sample density through the similarity of score vectors with neighboring data points, our method identifies underrepresented regions and upweights samples accordingly. This approach directly tackles inherent data imbalances, enhancing robustness by ensuring a more uniform dataset representation. Experiments on various tabular datasets under distribution shifts demonstrate that our method effectively improves performance on imbalanced data. Code is available at https://github.com/YunzeTong/ latent-score-based-reweighting .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng 等NeurIPS 2025 · 被引用 45 次
- Distributionally Robust Optimization via Generative Ambiguity ModelingJiaqi Wen, Jianyi YangICLR 2026 · 被引用 3 次
- CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language ModelsKairong Han, Wenshuo Zhao, Ziyu Zhao, Ye Jun Jian 等EMNLP 2025 · 被引用 3 次
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image GenerationZijing Hu, Yunze Tong, Fengda Zhang, Junkun Yuan 等ICLR 2026 · 被引用 3 次
- Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPOYunze Tong, Mushui Liu, Canyu Zhao, Wanggui He 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
相关 Paper
- Fairness with Adaptive WeightsJunyi Chai, Xiaoqian WangICML 2022 · 被引用 47 次
- Class-Conditional Distribution Balancing for Group Robust ClassificationMiaoyun Zhao, Qiang ZhangICML 2026 · 被引用 1 次
- Group-robust Sample Reweighting for Subpopulation Shifts via Influence FunctionsRui Qiao, Zhaoxuan Wu, Jingtan Wang, Pang Wei Koh 等ICLR 2025
- Software Fairness Dilemma: Is Bias Mitigation a Zero-Sum Game?Zhenpeng Chen, Xinyue Li, Jie M. Zhang, Weisong Sun 等FSE 2025
- Language-Interfaced Tabular Oversampling via Progressive Imputation and Self-AuthenticationJune Yong Yang, Geondo Park, Joowon Kim, Hyeongwon Jang 等ICLR 2024 · 被引用 8 次
