Counterfactual Residual Data Augmentation for Regression
Hossein Mohebbi, Oliver Schulte, Ke Li, Pascal Poupart
摘要
Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations. Inspired by the impact of data augmentation in vision and language, we propose a novel Counterfactual Residual Data Augmentation (CRDA) technique for tabular regression. Our key insight is that once a regressor has modeled the systematic component of the data, the remaining noise can be viewed as an invariant residual that remains stable under small perturbations of carefully selected features. We exploit this residual invariance to generate new, yet realistic, training samples, effectively expanding the dataset without requiring additional real data. Our method is model-agnostic and readily applicable to various types of regressors. In experiments across datasets from a variety of benchmark repositories, on average, CRDA reduces an MLP Regressor's MSE by 22.9% and an XGBoost Regressor's MSE by 6.4%. When compared to existing state-of-the-art data generators and augmentation techniques, CRDA consistently outperforms in MSE reduction. By adding principled counterfactual variations to the training data, our method offers a simple and efficient remedy for noise-prone, small-sample regression settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Cause-Effect Inference in Location-Scale Noise Models: Maximum Likelihood vs. Independence TestingXiangyu Sun, Oliver SchulteNeurIPS 2023 · 被引用 10 次
- Anchor Data AugmentationNora Schneider, Shirin Goshtasbpour, Fernando Pérez-CruzNeurIPS 2023 · 被引用 8 次
- When Shift Happens - Confounding Is to BlameAbbavaram Gowtham Reddy, Celia Rubio-Madrigal, Rebekka Burkholz, Krikamol MuandetICLR 2026 · 被引用 5 次
- An Analysis of Causal Effect Estimation using Outcome Invariant Data AugmentationUzair Akbar, Niki Kilbertus, Hao Shen, Krikamol Muandet 等NeurIPS 2025 · 被引用 3 次
相关 Paper
- APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-TuningHong-Wei Wu, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih PengAAAI 2025 · 被引用 1 次
- Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency GraphsShuo Yang, Zheyu Zhang, Bardh Prenkaj, Gjergji KasneciEMNLP 2025 · 被引用 6 次
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk 等ICLR 2023 · 被引用 45 次
- ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement LearningXiaofeng Lin, Seungbae Kim, Zhuoya Li, Zachary DeSoto 等ICML 2026
- Iterative Counterfactual Data AugmentationMitchell Plyler, Min ChiAAAI 2025 · 被引用 1 次
