Adaptive Data Augmentation for Supervised Learning over Missing Data
Tongyu Liu, Ju Fan, Yinqing Luo, Nan Tang, Guoliang Li, Xiaoyong Du
摘要
Real-world data is dirty, which causes serious problems in (supervised) machine learning (ML). The widely used practice in such scenario is to first repair the labeled source (a.k.a. train) data using rule-, statistical- or ML-based methods and then use the "repaired" source to train an ML model. During production, unlabeled target (a.k.a. test) data will also be repaired, and is then fed in the trained ML model for prediction. However, this process often causes a performance degradation when the source and target datasets are dirty with different noise patterns , which is common in practice.
In this paper, we propose an adaptive data augmentation approach, for handling missing data in supervised ML. The approach extracts noise patterns from target data, and adapts the source data with the extracted target noise patterns while still preserving supervision signals in the source. Then, it patches the ML model by retraining it on the adapted data, in order to better serve the target. To effectively support adaptive data augmentation, we propose a novel generative adversarial network (GAN) based framework, called DAGAN, which works in an unsupervised fashion. DAGAN consists of two connected GAN networks. The first GAN learns the noise pattern from the target, for target mask generation. The second GAN uses the learned target mask to augment the source data, for source data adaptation. The augmented source data is used to retrain the ML model. Extensive experiments show that our method significantly improves the ML model performance and is more robust than the state-of-the-art missing data imputation solutions for handling datasets with different missing value patterns.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Domain Adaptation for Deep Entity ResolutionJianhong Tu, Ju Fan, Nan Tang, Peng Wang 等SIGMOD 2022 · 被引用 46 次
- Efficient and Effective Data Imputation with Influence FunctionsXiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao 等VLDB 2022 · 被引用 38 次
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan 等SIGMOD 2023 · 被引用 37 次
- Identity-Disentangled Adversarial Augmentation for Self-supervised LearningKaiwen Yang, Tianyi Zhou, Xinmei Tian, Dacheng TaoICML 2022 · 被引用 13 次
- Controllable Tabular Data Synthesis Using Diffusion ModelsTongyu Liu, Ju Fan, Nan Tang, Guoliang Li 等SIGMOD 2024 · 被引用 13 次
它引用的顶会 Paper3
- Active Learning for ML Enhanced Database SystemsLin Ma, Bailu Ding, Sudipto Das, Adith SwaminathanSIGMOD 2020 · 被引用 57 次
- Baran: Effective Error Correction via a Unified Context Representation and Transfer LearningMohammad Mahdavi, Ziawasch AbedjanVLDB 2020
- Relational Data Synthesis using Generative Adversarial Networks: A Design Space ExplorationJu Fan, Tongyu Liu, Guoliang Li, Junyou Chen 等VLDB 2020
相关 Paper
- Meta-GAIN for Missing Data ImputationTao Tong, Xiaofeng Zhu, Jiangzhang GanAAAI 2026
- Improving the Training of the GANs with Limited Data via Dual Adaptive Noise InjectionZhaoyu Zhang, Yang Hua, Guanxiong Sun, Hui Wang 等ACM MM 2024 · 被引用 3 次
- Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language UnderstandingYingmei Guo, Linjun Shou, Jian Pei, Ming Gong 等EMNLP 2021 · 被引用 2 次
- Are labels informative in semi-supervised learning? Estimating and leveraging the missing-data mechanismAude Sportisse, Hugo Schmutz, Olivier Humbert, Charles Bouveyron 等ICML 2023 · 被引用 10 次
- Missing Data Imputation by Reducing Mutual Information with Rectified FlowsJiahao Yu, Qizhen Ying, Leyang Wang, Ziyue Jiang 等NeurIPS 2025 · 被引用 9 次
