inversedMixup: Data Augmentation via Inverting Mixed Embeddings
Fanshuang Kong, Richong Zhang, Qiyu Sun, Zhijie Nie, Ting Deng, Chunming Hu
Abstract
Mixup generates augmented samples by linearly interpolating inputs and labels with a controllable ratio. However, since it operates at the latent embedding level, the resulting samples are not human-interpretable. In contrast, LLM-based augmentation methods produce sentences via prompts at the token level, yielding readable outputs but offering limited control over the generation process. Inspired by recent advances in LLM inversion, which reconstructs natural language from embeddings and helps bridge the gap between latent embedding space and discrete token space, we propose inversedMixup, a unified framework that combines the controllability of Mixup with the interpretability of LLM-based generation. Specifically, inversedMixup aligns the output embedding space of a task-specific model with the input embedding space of an LLM, so that mixed embeddings can be reconstructed, under a controllable mixing ratio, into human-interpretable sentences. This interpretability provides the first empirical evidence of the manifold intrusion phenomenon in text Mixup. Building on this, we extend inversedMixup into a three-stage data augmentation method, and introduce a simple yet effective strategy to mitigate manifold intrusion during augmentation. Extensive experiments demonstrate the effectiveness and generalizability of our approach in both few-shot and fully supervised scenarios. Our code is available at https://github.com/pypi1412/inversedMixup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7c9cf08-a27a-4633-831c-185120b0aa2aBuilds on12
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Adversarial Domain Adaptation with Domain MixupMinghao Xu, Jian Zhang, Bingbing Ni, Teng Li et al.AAAI 2020 · 499 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- Nonlinear Mixup: Out-Of-Manifold Data Augmentation for Text ClassificationHongyu GuoAAAI 2020 · 124 citations
- Mixup Inference: Better Exploiting Mixup to Defend Adversarial AttacksTianyu Pang, Kun Xu, Jun ZhuICLR 2020 · 114 citations
Related papers
- PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationGaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, Issam H. LaradjiEMNLP 2023 · 11 citations
- OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio SeparationTanvir Mahmud, Diana MarculescuEMNLP 2024 · 1 citation
- TextManiA: Enriching Visual Feature by Text-driven Manifold AugmentationMoon Ye-Bin, Jisoo Kim, Hongyeob Kim, Kilho Son et al.ICCV 2023 · 14 citations
- DiffusionCLIP: Text-Guided Diffusion Models for Robust Image ManipulationGwanghyun Kim, Taesung Kwon, Jong Chul YeCVPR 2022 · 458 citations
- Reverse Prompt Engineering: A Zero-Shot, Genetic Algorithm Approach to Language Model InversionHanqing Li, Diego KlabjanEMNLP 2025 · 1 citation
