Distilling an End-to-End Voice Assistant Without Instruction Training Data
William Barr Held, Yanzhe Zhang, Weiyan Shi, Minzhi Li, Michael J. Ryan, Diyi Yang
摘要
Voice assistants, such as Siri and Google Assistant, typically model audio and text separately, resulting in lost speech information and increased complexity. Recent efforts to address this with end-to-end Speech Large Language Models (speech-in, text-out) trained with supervised finetuning (SFT) have led to models "forgetting" capabilities from text-only LLMs. Our work proposes an alternative paradigm for training Speech LLMs without instruction data, using the response of a text-only LLM to transcripts as self-supervision. Importantly, this process can be performed without annotated responses. We show that our Distilled Voice Assistant (DiVA) generalizes to Spoken Question Answering, Classification, and Translation. Furthermore, DiVA better matches user preferences, achieving a 72% win rate compared with state-of-the-art models like Qwen 2 Audio, despite using >100x less training compute.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language ModelsZirui Song, Qian Jiang, Mingxuan Cui, Mingzhe Li 等ACL 2026 · 被引用 20 次
- DIFFA: Large Language Diffusion Models Can Listen and UnderstandJiaming Zhou, Hongjie Chen, Shiwan Zhao, Jian Kang 等AAAI 2026 · 被引用 10 次
- SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal SteeringWeilin Lin, Jianze Li, Hui Xiong, Li LiuICML 2026 · 被引用 6 次
- Mind the Gap: Static and Interactive Evaluations of Large Audio ModelsMinzhi Li, William Barr Held, Michael J. Ryan, Kunat Pipatanakul 等ACL 2025
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 等ICLR 2024 · 被引用 557 次
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 被引用 488 次
相关 Paper
- Soundwave: Less is More for Speech-Text Alignment in LLMsYuhao Zhang, Zhiheng Liu, Fan Bu, Ruiyu Zhang 等ACL 2025 · 被引用 11 次
- Towards True Speech-to-Speech Models Without Text GuidanceXingjian Zhao, Zhe Xu, Luozhijie Jin, Yang Wang 等ICLR 2026 · 被引用 8 次
- Recent Advances in Speech Language Models: A SurveyWenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng 等ACL 2025
- Exploring Transfer Learning For End-to-End Spoken Language UnderstandingSubendhu Rongali, Beiye Liu, Liwei Cai, Konstantine Arkoudas 等AAAI 2021 · 被引用 26 次
- Introducing Semantics into Speech EncodersDerek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim 等ACL 2023 · 被引用 2 次
