End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering
Jiliang Hu, Zuchao Li, Baoyuan Qi, Guoming Liu, Ping Wang
摘要
Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help preprocessing long-form speech. But the performance of existing speech-related retrievers is lacking. To address this challenge, we propose CLSR, an end-to-end contrastive language-speech retriever that efficiently extracts question-relevant segments from long audio recordings for downstream SQA task. Unlike conventional speech-text contrastive models, CLSR incorporates an intermediate step that converts acoustic features into text-like representations prior to alignment, thereby more effectively bridging the gap between modalities. Experimental results across four cross-modal retrieval datasets demonstrate that CLSR surpasses both end-to-end speech related retrievers and pipeline approaches combining speech recognition with text retrieval, providing a robust foundation for advancing practical long-form SQA applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- RobuTrans: A Robust Transformer-Based Text-to-Speech ModelNaihan Li, Yanqing Liu, Yu Wu, Shujie Liu 等AAAI 2020 · 被引用 43 次
- SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding TasksSuwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad 等ACL 2023 · 被引用 21 次
相关 Paper
- LSAR: Sparse Lexical Representation Learning for Efficient and Interpretable Audio RetrievalHaoyue Li, Yuzhe Bai, Li NiuKDD 2026
- WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue ModelsYifu Chen, Shengpeng Ji, Haoxiao Wang, Ziqing Wang 等ACL 2025
- Multimodal Hypothetical Summary for Retrieval-based Multi-image Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin 等AAAI 2025 · 被引用 1 次
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 被引用 12 次
- Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature AlignmentSarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg 等ICCV 2023 · 被引用 43 次
