InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
Dingdong Wang, Jin Xu, Ruihang Chu, Zhifang Guo, Xiong Wang, Jincenzi Wu, Dongchao Yang, Shengpeng Ji, Junyang Lin
摘要
Recent advancements in speech large language models (SpeechLLMs) have attracted considerable attention. Nonetheless, current methods exhibit suboptimal performance in adhering to speech instructions. Notably, the intelligence of models significantly diminishes when processing speech-form input as compared to direct text-form input. Prior work has attempted to mitigate this semantic inconsistency between speech and text representations through techniques such as representation and behavior alignment, which involve the meticulous design of data pairs during the post-training phase. In this paper, we introduce a simple and scalable training method called InSerter, which stands for Interleaved Speech-Text Representation Pre-training. InSerter is designed to pre-train large-scale unsupervised speech-text sequences, where the speech is synthesized from randomly selected segments of an extensive text corpus using text-to-speech conversion. Consequently, the model acquires the ability to generate textual continuations corresponding to the provided speech segments, obviating the need for intensive data design endeavors. To systematically evaluate speech instruction-following capabilities, we introduce SpeechInstructBench, the first comprehensive benchmark specifically designed for speech-oriented instruction-following tasks. Our proposed model InSerter achieves SOTA performance in SpeechInstructBench and demonstrates superior or competitive results across diverse speech processing tasks. SpeechInstructBench is available at https://huggingface.co/datasets/ ddwang2000/SpeechInstructBench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang 等ICLR 2026 · 被引用 143 次
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific TalksSara Papi, Maike Züfle, Marco Gaido, Beatrice Savoldi 等ICLR 2026 · 被引用 20 次
- Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMsDingdong Wang, Junan Li, Mingyu Cui, Dongchao Yang 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 等ICLR 2024 · 被引用 557 次
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong 等ACL 2024 · 被引用 10 次
- Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue AgentsBandhav Veluri, Benjamin N. Peloquin, Bokai Yu, Hongyu Gong 等EMNLP 2024 · 被引用 8 次
- Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language ModelsKuofeng Gao, Shutao Xia, Ke Xu, Philip Torr 等ACL 2025
相关 Paper
- SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-TuningPrabhat Pandey, Rupak Vignesh Swaminathan, K. V. Vijay Girish, Arunasish Sen 等ACL 2025 · 被引用 10 次
- Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningCongrui Du, Yang Zhang, Kaizhi Qian, Shiyu ChangICML 2026
- InstructSpeech: Following Speech Editing Instructions via Large Language ModelsRongjie Huang, Ruofan Hu, Yongqi Wang, Zehan Wang 等ICML 2024 · 被引用 10 次
- Self-Powered LLM Modality Expansion for Large Speech-Text ModelsTengfei Yu, Xuebo Liu, Zhiyi Hou, Liang Ding 等EMNLP 2024 · 被引用 1 次
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech ModelsFeng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue 等ACL 2026 · 被引用 17 次
