Quality-Aware Self-Training on Differentiable Synthesis of Rare Relational Data
Chongsheng Zhang, Yaxin Hou, Ke Chen, Shuang Cao, Gaojuan Fan, Ji Liu
摘要
Data scarcity is a very common real-world problem that poses a major challenge to data-driven analytics. Although a lot of data-balancing approaches have been proposed to mitigate this problem, they may drop some useful information or fall into the overfitting problem. Generative Adversarial Network (GAN) based data synthesis methods can alleviate such a problem but lack of quality control over the generated samples. Moreover, the latent associations between the attribute set and the class labels in a relational data cannot be easily captured by a vanilla GAN. In light of this, we introduce an end-to-end self-training scheme (namely, Quality-Aware Self-Training) for rare relational data synthesis, which generates labeled synthetic data via pseudo labeling on GAN-based synthesis. We design a semantic pseudo labeling module to first control the quality of the generated features/samples, then calibrate their semantic labels via a classifier committee consisting of multiple pre-trained shallow classifiers. The high-confident generated samples with calibrated pseudo labels are then fed into a semantic classification network as augmented samples for self-training. We conduct extensive experiments on 20 benchmark datasets of different domains, including 14 industrial datasets. The results show that our method significantly outperforms state-of-the-art methods, including two recent GAN-based data synthesis schemes. Codes are available at https://github.com/yaxinhou/QAST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Keep It on a Leash: Controllable Pseudo-label Generation Towards Realistic Long-Tailed Semi-Supervised LearningYaxin Hou, Bo Han, Yuheng Jia, Hui Liu 等NeurIPS 2025 · 被引用 4 次
- A Square Peg in a Square Hole: Meta-Expert for Long-Tailed Semi-Supervised LearningYaxin Hou, Yuheng JiaICML 2025
- Beyond Distribution Estimation: Simplex Anchored Structural Inference Towards Universal Semi-Supervised LearningYaxin Hou, Jun Ma, Hanyang Li, Bo Han 等ICML 2026
它引用的顶会 Paper7
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui 等NeurIPS 2020 · 被引用 755 次
- ST++: Make Self-trainingWork Better for Semi-supervised Semantic SegmentationLihe Yang, Wei Zhuo, Lei Qi, Yinghuan Shi 等CVPR 2022 · 被引用 467 次
- Projected GANs Converge FasterAxel Sauer, Kashyap Chitta, Jens Müller, Andreas GeigerNeurIPS 2021 · 被引用 325 次
- The Majority Can Help the Minority: Context-rich Minority Oversampling for Long-tailed ClassificationSeulki Park, Youngkyu Hong, Byeongho Heo, Sangdoo Yun 等CVPR 2022 · 被引用 199 次
- Targeted Supervised Contrastive Learning for Long-Tailed RecognitionTianhong Li, Peng Cao, Yuan Yuan, Lijie Fan 等CVPR 2022 · 被引用 196 次
相关 Paper
- Can Pseudo-Label Be More Reliable? A Simple yet Effective Topology-Aware Graph Self-Training MethodGen Liu, Zhongying Zhao, Hui Zhou, Chao Li 等AAAI 2026
- Relational Data Synthesis using Generative Adversarial Networks: A Design Space ExplorationJu Fan, Tongyu Liu, Guoliang Li, Junyou Chen 等VLDB 2020
- A Unified Generative Adversarial Network Training via Self-Labeling and Self-AttentionTomoki Watanabe, Paolo FavaroICML 2021 · 被引用 3 次
- RareGAN: Generating Samples for Rare ClassesZinan Lin, Hao Liang, Giulia Fanti, Vyas SekarAAAI 2022 · 被引用 14 次
- Graph Contrastive Learning with Generative Adversarial NetworkCheng Wu, Chaokun Wang, Jingcao Xu, Ziyang Liu 等KDD 2023 · 被引用 33 次
