inversedMixup: Data Augmentation via Inverting Mixed Embeddings
Fanshuang Kong, Richong Zhang, Qiyu Sun, Zhijie Nie, Ting Deng, Chunming Hu
摘要
Mixup generates augmented samples by linearly interpolating inputs and labels with a controllable ratio. However, since it operates at the latent embedding level, the resulting samples are not human-interpretable. In contrast, LLM-based augmentation methods produce sentences via prompts at the token level, yielding readable outputs but offering limited control over the generation process. Inspired by recent advances in LLM inversion, which reconstructs natural language from embeddings and helps bridge the gap between latent embedding space and discrete token space, we propose inversedMixup, a unified framework that combines the controllability of Mixup with the interpretability of LLM-based generation. Specifically, inversedMixup aligns the output embedding space of a task-specific model with the input embedding space of an LLM, so that mixed embeddings can be reconstructed, under a controllable mixing ratio, into human-interpretable sentences. This interpretability provides the first empirical evidence of the manifold intrusion phenomenon in text Mixup. Building on this, we extend inversedMixup into a three-stage data augmentation method, and introduce a simple yet effective strategy to mitigate manifold intrusion during augmentation. Extensive experiments demonstrate the effectiveness and generalizability of our approach in both few-shot and fully supervised scenarios. Our code is available at https://github.com/pypi1412/inversedMixup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Adversarial Domain Adaptation with Domain MixupMinghao Xu, Jian Zhang, Bingbing Ni, Teng Li 等AAAI 2020 · 被引用 499 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Nonlinear Mixup: Out-Of-Manifold Data Augmentation for Text ClassificationHongyu GuoAAAI 2020 · 被引用 124 次
- Mixup Inference: Better Exploiting Mixup to Defend Adversarial AttacksTianyu Pang, Kun Xu, Jun ZhuICLR 2020 · 被引用 114 次
相关 Paper
- PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationGaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, Issam H. LaradjiEMNLP 2023 · 被引用 11 次
- OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio SeparationTanvir Mahmud, Diana MarculescuEMNLP 2024 · 被引用 1 次
- TextManiA: Enriching Visual Feature by Text-driven Manifold AugmentationMoon Ye-Bin, Jisoo Kim, Hongyeob Kim, Kilho Son 等ICCV 2023 · 被引用 14 次
- DiffusionCLIP: Text-Guided Diffusion Models for Robust Image ManipulationGwanghyun Kim, Taesung Kwon, Jong Chul YeCVPR 2022 · 被引用 458 次
- Reverse Prompt Engineering: A Zero-Shot, Genetic Algorithm Approach to Language Model InversionHanqing Li, Diego KlabjanEMNLP 2025 · 被引用 1 次
