Adaptive Data Augmentation for Supervised Learning over Missing Data
Tongyu Liu, Ju Fan, Yinqing Luo, Nan Tang, Guoliang Li, Xiaoyong Du
Abstract
Real-world data is dirty, which causes serious problems in (supervised) machine learning (ML). The widely used practice in such scenario is to first repair the labeled source (a.k.a. train) data using rule-, statistical- or ML-based methods and then use the "repaired" source to train an ML model. During production, unlabeled target (a.k.a. test) data will also be repaired, and is then fed in the trained ML model for prediction. However, this process often causes a performance degradation when the source and target datasets are dirty with different noise patterns , which is common in practice.
In this paper, we propose an adaptive data augmentation approach, for handling missing data in supervised ML. The approach extracts noise patterns from target data, and adapts the source data with the extracted target noise patterns while still preserving supervision signals in the source. Then, it patches the ML model by retraining it on the adapted data, in order to better serve the target. To effectively support adaptive data augmentation, we propose a novel generative adversarial network (GAN) based framework, called DAGAN, which works in an unsupervised fashion. DAGAN consists of two connected GAN networks. The first GAN learns the noise pattern from the target, for target mask generation. The second GAN uses the learned target mask to augment the source data, for source data adaptation. The augmented source data is used to retrain the ML model. Extensive experiments show that our method significantly improves the ML model performance and is more robust than the state-of-the-art missing data imputation solutions for handling datasets with different missing value patterns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66ad2bab-6ba2-4a6d-9de4-a4597b74f085Cited by top-tier papers11
- Domain Adaptation for Deep Entity ResolutionJianhong Tu, Ju Fan, Nan Tang, Peng Wang et al.SIGMOD 2022 · 46 citations
- Efficient and Effective Data Imputation with Influence FunctionsXiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao et al.VLDB 2022 · 38 citations
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan et al.SIGMOD 2023 · 37 citations
- Identity-Disentangled Adversarial Augmentation for Self-supervised LearningKaiwen Yang, Tianyi Zhou, Xinmei Tian, Dacheng TaoICML 2022 · 13 citations
- Controllable Tabular Data Synthesis Using Diffusion ModelsTongyu Liu, Ju Fan, Nan Tang, Guoliang Li et al.SIGMOD 2024 · 13 citations
Builds on3
- Active Learning for ML Enhanced Database SystemsLin Ma, Bailu Ding, Sudipto Das, Adith SwaminathanSIGMOD 2020 · 57 citations
- Baran: Effective Error Correction via a Unified Context Representation and Transfer LearningMohammad Mahdavi, Ziawasch AbedjanVLDB 2020
- Relational Data Synthesis using Generative Adversarial Networks: A Design Space ExplorationJu Fan, Tongyu Liu, Guoliang Li, Junyou Chen et al.VLDB 2020
Related papers
- Meta-GAIN for Missing Data ImputationTao Tong, Xiaofeng Zhu, Jiangzhang GanAAAI 2026
- Improving the Training of the GANs with Limited Data via Dual Adaptive Noise InjectionZhaoyu Zhang, Yang Hua, Guanxiong Sun, Hui Wang et al.ACM MM 2024 · 3 citations
- Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language UnderstandingYingmei Guo, Linjun Shou, Jian Pei, Ming Gong et al.EMNLP 2021 · 2 citations
- Are labels informative in semi-supervised learning? Estimating and leveraging the missing-data mechanismAude Sportisse, Hugo Schmutz, Olivier Humbert, Charles Bouveyron et al.ICML 2023 · 10 citations
- Missing Data Imputation by Reducing Mutual Information with Rectified FlowsJiahao Yu, Qizhen Ying, Leyang Wang, Ziyue Jiang et al.NeurIPS 2025 · 9 citations
