Back-Modality: Leveraging Modal Transformation for Data Augmentation
Zhi Li, Yifan Liu, Yin Zhang
摘要
We introduce Back-Modality, a novel data augmentation schema predicated on modal transformation. Data from an initial modality undergo a transformation to an intermediate modality, followed by a reverse transformation. This framework serves dual roles. On one hand, it operates as a general data augmentation strategy. On the other hand, it allows for other augmentation techniques, suitable for the intermediate modality, to enhance the initial modality. For instance, data augmentation methods applicable to pure text can be employed to augment images, thereby facilitating the cross-modality of data augmentation techniques. To validate the viability and efficacy of our framework, we proffer three instantiations of Back-Modality: back-captioning, back-imagination, and back-speech. Comprehensive evaluations across tasks such as image classification, sentiment classification, and textual entailment demonstrate that our methods consistently enhance performance under data-scarce circumstances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
相关 Paper
- RAG4DMC: Retrieval-Augmented Generation for Data-Level Modality CompletionNingxin He, Yongheng Deng, Sheng Yue, Yongjian Fu 等ICLR 2026
- MODALS: Modality-agnostic Automated Data Augmentation in the Latent SpaceTsz-Him Cheung, Dit-Yan YeungICLR 2021 · 被引用 22 次
- Unpaired Image-to-Speech Synthesis With Multimodal Information BottleneckShuang Ma, Daniel McDuff, Yale SongICCV 2019 · 被引用 29 次
- Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language LearningShivaen Ramshetty, Gaurav Verma, Srijan KumarACL 2023 · 被引用 4 次
- TMMDA: A New Token Mixup Multimodal Data Augmentation for Multimodal Sentiment AnalysisXianbing Zhao, Yixin Chen, Sicen Liu, Xuan Zang 等WWW 2023 · 被引用 17 次
