Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular Data
Yunze Tong, Fengda Zhang, Zihao Tang, Kaifeng Gao, Kai Huang, Pengfei Lyu, Jun Xiao, Kun Kuang
Abstract
Machine learning models often perform well on tabular data by optimizing average prediction accuracy. However, they may underperform on specific subsets due to inherent biases in the training data, such as associations with noncausal features like demographic information. These biases lead to critical robustness issues as models may inherit or amplify them, resulting in poor performance where such misleading correlations do not hold. Existing mitigation methods have significant limitations: some require prior group labels, which are often unavailable, while others focus solely on the conditional distribution P (Y |X), upweighting misclassified samples without effectively balancing the overall data distribution P (X). To address these shortcomings, we propose a latent score-based reweighting framework. It leverages score-based models to capture the joint data distribution P (X, Y ) without relying on additional prior information. By estimating sample density through the similarity of score vectors with neighboring data points, our method identifies underrepresented regions and upweights samples accordingly. This approach directly tackles inherent data imbalances, enhancing robustness by ensuring a more uniform dataset representation. Experiments on various tabular datasets under distribution shifts demonstrate that our method effectively improves performance on imbalanced data. Code is available at https://github.com/YunzeTong/ latent-score-based-reweighting .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng et al.NeurIPS 2025 · 45 citations
- Distributionally Robust Optimization via Generative Ambiguity ModelingJiaqi Wen, Jianyi YangICLR 2026 · 3 citations
- CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language ModelsKairong Han, Wenshuo Zhao, Ziyu Zhao, Ye Jun Jian et al.EMNLP 2025 · 3 citations
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image GenerationZijing Hu, Yunze Tong, Fengda Zhang, Junkun Yuan et al.ICLR 2026 · 3 citations
- Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPOYunze Tong, Mushui Liu, Canyu Zhao, Wanggui He et al.ICML 2026 · 3 citations
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Fairness with Adaptive WeightsJunyi Chai, Xiaoqian WangICML 2022 · 47 citations
- Class-Conditional Distribution Balancing for Group Robust ClassificationMiaoyun Zhao, Qiang ZhangICML 2026 · 1 citation
- Group-robust Sample Reweighting for Subpopulation Shifts via Influence FunctionsRui Qiao, Zhaoxuan Wu, Jingtan Wang, Pang Wei Koh et al.ICLR 2025
- Software Fairness Dilemma: Is Bias Mitigation a Zero-Sum Game?Zhenpeng Chen, Xinyue Li, Jie M. Zhang, Weisong Sun et al.FSE 2025
- Language-Interfaced Tabular Oversampling via Progressive Imputation and Self-AuthenticationJune Yong Yang, Geondo Park, Joowon Kim, Hyeongwon Jang et al.ICLR 2024 · 8 citations
