Speech-Text Pre-training for Spoken Dialog Understanding with Explicit Cross-Modal Alignment
Tianshu Yu, Haoyu Gao, Ting-En Lin, Min Yang, Yuchuan Wu, Wentao Ma, Chao Wang, Fei Huang, Yongbin Li
摘要
Recently, speech-text pre-training methods have shown remarkable success in many speech and natural language processing tasks. However, most previous pre-trained models are usually tailored for one or two specific tasks, but fail to conquer a wide range of speech-text tasks. In addition, existing speech-text pretraining methods fail to explore the contextual information within a dialogue to enrich utterance representations. In this paper, we propose Speech-text dialog Pre-training for spoken dialog understanding with ExpliCiT cRoss-Modal Alignment (SPECTRA), which is the first-ever speech-text dialog pre-training model. Concretely, to consider the temporality of speech modality, we design a novel temporal position prediction task to capture the speech-text alignment. This pre-training task aims to predict the start and end time of each textual word in the corresponding speech waveform. In addition, to learn the characteristics of spoken dialogs, we generalize a response selection task from textual dialog pre-training to speech-text dialog pre-training scenarios. Experimental results on four different downstream speech-text tasks demonstrate the superiority of SPECTRA in learning speech-text alignment and multi-turn dialog context. 1 * Equal contribution. This work was conducted when Tianshu Yu and Haoyu Gao were interning at Alibaba.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent RecognitionQianrui Zhou, Hua Xu, Hao Li, Hanlei Zhang 等AAAI 2024 · 被引用 45 次
- UniSA: Unified Generative Framework for Sentiment AnalysisZaijing Li, Ting-En Lin, Yuchuan Wu, Meng Liu 等ACM MM 2023 · 被引用 22 次
- EmoWear: Exploring Emotional Teasers for Voice Message Interaction on SmartwatchesPengcheng An, Jiawen Stefanie Zhu, Zibo Zhang, Yifei Yin 等CHI 2024 · 被引用 19 次
- Semi-IIN: Semi-Supervised Intra-Inter Modal Interaction Learning Network for Multimodal Sentiment AnalysisJinhao Lin, Yifei Wang, Yanwu Xu, Qi LiuAAAI 2025 · 被引用 3 次
- Proactive Hearing Assistants that Isolate Egocentric ConversationsGuilin Hu, Malek Itani, Tuochao Chen, Shyamnath GollakotaEMNLP 2025 · 被引用 1 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh 等ACL 2020 · 被引用 584 次
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu 等ACL 2020 · 被引用 229 次
相关 Paper
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun 等SIGIR 2022 · 被引用 41 次
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu 等EMNLP 2022 · 被引用 38 次
- FutureTOD: Teaching Future Knowledge to Pre-trained Language Model for Task-Oriented DialogueWeihao Zeng, Keqing He, Yejie Wang, Chen Zeng 等ACL 2023 · 被引用 3 次
- HiTeA: Hierarchical Temporal-Aware Video-Language Pre-trainingQinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu 等ICCV 2023 · 被引用 102 次
- A3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and EditingHe Bai, Renjie Zheng, Jun-Kun Chen, Mingbo Ma 等ICML 2022 · 被引用 64 次
