Self-Supervised Audio-and-Text Pre-training with Extremely Low-Resource Parallel Data
Yu Kang, Tianqiao Liu, Hang Li, Yang Hao, Wenbiao Ding
Abstract
Multimodal pre-training for audio-and-text has recently been proved to be effective and has significantly improved the performance of many downstream speech understanding tasks. However, these state-of-the-art pre-training audio-text models work well only when provided with large amount of parallel audio-and-text data, which brings challenges on many languages that are rich in unimodal corpora but scarce of parallel cross-modal corpus. In this paper, we investigate whether it is possible to pre-train an audio-text multimodal model with extremely low-resource parallel data and extra non-parallel unimodal data. Our pre-training framework consists of the following components: (1) Intra-modal Denoising Auto-Encoding (IDAE), which is able to reconstruct input text (audio) representations from a noisy version of itself. (2) Cross-modal Denoising Auto-Encoding (CDAE), which is pre-trained to reconstruct the input text (audio), given both a noisy version of the input text (audio) and the corresponding translated noisy audio features (text embeddings). ( 3 ) Iterative Denoising Process (IDP), which iteratively translates raw audio (text) and the corresponding text embeddings (audio features) translated from previous iteration into the new less-noisy text embeddings (audio features). We adapt a dual cross-modal Transformer as our backbone model which consists of two unimodal encoders for IDAE and two cross-modal encoders for CDAE and IDP. Our method achieves comparable performance on multiple downstream speech understanding tasks compared with the model pre-trained on fully parallel data, demonstrating the great potential of the proposed method. Our code is available at: https://github.com/KarlYuKang/Low-Resource-Multimodal-Pre-training .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24c6703a-9f69-4622-9d84-2a74504a6961Cited by top-tier papers4
- Speech-Text Pre-training for Spoken Dialog Understanding with Explicit Cross-Modal AlignmentTianshu Yu, Haoyu Gao, Ting-En Lin, Min Yang et al.ACL 2023 · 26 citations
- From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint TrainingTianqiao Liu, Xueyi Li, Hao Wang, Haoxuan Li et al.ICLR 2026 · 6 citations
- EM-Network: Oracle Guided Self-distillation for Sequence LearningJi Won Yoon, Sunghwan Ahn, Hyeonseung Lee, Minchan Kim et al.ICML 2023 · 3 citations
- Heuristic-free Knowledge Distillation for Streaming ASR via Multi-modal TrainingJi Won YoonAAAI 2025
Builds on3
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- CTAL: Pre-training Cross-modal Transformer for Audio-and-Language RepresentationsHang Li, Wenbiao Ding, Yu Kang, Tianqiao Liu et al.EMNLP 2021 · 9 citations
Related papers
- BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic DataXuenan Xu, Zhiling Zhang, Zelin Zhou, Pingyue Zhang et al.ACM MM 2023 · 12 citations
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang et al.AAAI 2023 · 17 citations
- UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task LearnerDongchao Yang, Haohan Guo, Yuanyuan Wang, Rongjie Huang et al.NeurIPS 2024 · 55 citations
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang et al.ACM MM 2022 · 26 citations
- NLIP: Noise-Robust Language-Image Pre-trainingRunhui Huang, Yanxin Long, Jianhua Han, Hang Xu et al.AAAI 2023 · 44 citations
