DiffAug: Enhance Unsupervised Contrastive Learning with Domain-Knowledge-Free Diffusion-based Data Augmentation
Zelin Zang, Hao Luo, Kai Wang, Panpan Zhang, Fan Wang, Stan Z. Li, Yang You
摘要
Unsupervised Contrastive learning has gained prominence in fields such as vision, and biology, leveraging predefined positive/negative samples for representation learning. Data augmentation, categorized into hand-designed and model-based methods, has been identified as a crucial component for enhancing contrastive learning. However, hand-designed methods require human expertise in domain-specific data while sometimes distorting the meaning of the data. In contrast, generative model-based approaches usually require supervised or large-scale external data, which has become a bottleneck constraining model training in many domains. To address the problems presented above, this paper proposes DiffAug, a novel unsupervised contrastive learning technique with diffusion mode-based positive data generation. DiffAug consists of a semantic encoder and a conditional diffusion model; the conditional diffusion model generates new positive samples conditioned on the semantic encoding to serve the training of unsupervised contrast learning. With the help of iterative training of the semantic encoder and diffusion model, DiffAug improves the representation ability in an uninterrupted and unsupervised manner. Experimental evaluations show that DiffAug outperforms hand-designed and SOTA model-based augmentation methods on DNA sequence, visual, and bio-feature datasets. The code for review is released at https://github.com/zangzelin/code_diffaug.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PhyloGen: Language Model-Enhanced Phylogenetic Inference via Graph Structure GenerationChenrui Duan, Zelin Zang, Siyuan Li, Yongjie Xu 等NeurIPS 2024 · 被引用 8 次
- Can Generative Models Improve Self-Supervised Representation Learning?Sana Ayromlou, Vahid Reza Khazaie, Fereshteh Forghani, Arash AfkanpourAAAI 2025 · 被引用 5 次
- How does Labeling Error Impact Contrastive Learning? A Perspective from Data Dimensionality ReductionJun Chen, Hong Chen, Yonghua Yu, Yiming YingICML 2025
它引用的顶会 Paper27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Do Generated Data Always Help Contrastive Learning?Yifei Wang, Jizhe Zhang, Yisen WangICLR 2024 · 被引用 36 次
- Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image GenerationGuojin Zhong, Jin Yuan, Pan Wang, Kailun Yang 等ACM MM 2023 · 被引用 7 次
- Adversarial Contrastive Graph Augmentation with Counterfactual RegularizationTao Long, Lei Zhang, Liang Zhang, Laizhong CuiAAAI 2025 · 被引用 5 次
- D2C: Diffusion-Decoding Models for Few-Shot Conditional GenerationAbhishek Sinha, Jiaming Song, Chenlin Meng, Stefano ErmonNeurIPS 2021 · 被引用 149 次
- RLEG: Vision-Language Representation Learning with Diffusion-based Embedding GenerationLiming Zhao, Kecheng Zheng, Yun Zheng, Deli Zhao 等ICML 2023 · 被引用 11 次
