Back-Modality: Leveraging Modal Transformation for Data Augmentation
Zhi Li, Yifan Liu, Yin Zhang
Abstract
We introduce Back-Modality, a novel data augmentation schema predicated on modal transformation. Data from an initial modality undergo a transformation to an intermediate modality, followed by a reverse transformation. This framework serves dual roles. On one hand, it operates as a general data augmentation strategy. On the other hand, it allows for other augmentation techniques, suitable for the intermediate modality, to enhance the initial modality. For instance, data augmentation methods applicable to pure text can be employed to augment images, thereby facilitating the cross-modality of data augmentation techniques. To validate the viability and efficacy of our framework, we proffer three instantiations of Back-Modality: back-captioning, back-imagination, and back-speech. Comprehensive evaluations across tasks such as image classification, sentiment classification, and textual entailment demonstrate that our methods consistently enhance performance under data-scarce circumstances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
Related papers
- RAG4DMC: Retrieval-Augmented Generation for Data-Level Modality CompletionNingxin He, Yongheng Deng, Sheng Yue, Yongjian Fu et al.ICLR 2026
- MODALS: Modality-agnostic Automated Data Augmentation in the Latent SpaceTsz-Him Cheung, Dit-Yan YeungICLR 2021 · 22 citations
- Unpaired Image-to-Speech Synthesis With Multimodal Information BottleneckShuang Ma, Daniel McDuff, Yale SongICCV 2019 · 29 citations
- Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language LearningShivaen Ramshetty, Gaurav Verma, Srijan KumarACL 2023 · 4 citations
- TMMDA: A New Token Mixup Multimodal Data Augmentation for Multimodal Sentiment AnalysisXianbing Zhao, Yixin Chen, Sicen Liu, Xuan Zang et al.WWW 2023 · 17 citations
