Domain Gap Embeddings for Generative Dataset Augmentation
Yinong Oliver Wang, Younjoon Chung, Chen Henry Wu, Fernando De la Torre
摘要
The performance of deep learning models is intrinsically tied to the quality, volume, and relevance of their training data. Gathering ample data for production scenarios of-ten demands significant time and resources. Among various strategies, data augmentation circumvents exhaustive data collection by generating new data points from existing ones. However, traditional augmentation techniques can be less effective amidst a shift in training and testing distributions. This paper explores the potential of synthetic data by leveraging large pre-trained models for data augmentation, especially when confronted with distribution shifts. Al-though recent advancements in generative models have en-abled several prior works in cross-distribution data gener-ation, they require model fine-tuning and a complex setup. To bypass these shortcomings, we introduce Domain Gap Embeddings (DoGE), a plug-and-play semantic data aug-mentation framework in a cross-distribution few-shot set-ting. Our method extracts disparities between source and desired data distributions in a latent form, and subsequently steers a generative process to supplement the training set with endless diverse synthetic samples. Our evaluations, conducted on a subpopulation shift and three domain adap-tation scenarios under afew-shot paradigm, reveal that our versatile method improves performance across tasks with-out needing hands-on intervention or intricate fine-tuning. DoGE paves the way to effortlessly generate realistic, con-trollable synthetic datasets following the test distributions, bolstering real-world efficacy for downstream task models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Generating Informative Samples for Risk-Averse Fine-Tuning of Downstream TasksHeasung Kim, Taekyun Lee, Hyeji Kim, Gustavo de VecianaNeurIPS 2025 · 被引用 2 次
- Improving Text-to-Image Generation with Intrinsic Self-Confidence RewardsSeungwook Kim, Minsu ChoCVPR 2026 · 被引用 1 次
- Focus On This, Not That! Steering LLMs with Adaptive Feature SpecificationTom A. Lamb, Adam Davies, Alasdair Paren, Philip Torr 等ICML 2025
- Rethinking Bias in Generative Data Augmentation for Medical AI: A Frequency Recalibration MethodChi Liu, Jincheng Liu, Congcong Zhu, Minghao Wang 等AAAI 2026
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak SupervisionXiao Fang, Minhyek Jeon, Shuowen Hu, Zheyang Qin 等ICCV 2025
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Effective Data Augmentation With Diffusion ModelsBrandon Trabucco, Kyle Doherty, Max Gurinas, Ruslan SalakhutdinovICLR 2024 · 被引用 380 次
- Few-shot Adaptation to Distribution Shifts By Mixing Source and Target EmbeddingsYihao Xue, Ali Payani, Yu Yang, Baharan MirzasoleimanICML 2024 · 被引用 4 次
- STraTA: Self-Training with Task Augmentation for Better Few-shot LearningTu Vu, Minh-Thang Luong, Quoc V. Le, Grady Simon 等EMNLP 2021 · 被引用 25 次
- Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot LearningYu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang 等ICML 2023 · 被引用 64 次
- OntoAug: Rethinking Generative Data Augmentation via Ontology GuidanceShuo Wang, Zhichuan Wang, Jun LuoCVPR 2026
