Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
Eliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar, Chulayuth Asawaroengchai, Soroosh Mariooryad, Ehud Rivlin, R. J. Skerry-Ryan, Michelle Tadmor Ramanovich
Abstract
We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire system is trained end-to-end and operates directly on spectrograms, simplifying our architecture. Key to our approach is a training objective that jointly supervises speech recognition, text continuation, and speech synthesis using only paired speech-text pairs, enabling a `cross-modal' chain-of-thought within a single decoding pass. Our method surpasses existing spoken language models in speaker preservation and semantic coherence. Furthermore, the proposed model improves upon direct initialization in retaining the knowledge of the original LLM as demonstrated through spoken QA datasets. We release our audio samples (https://michelleramanovich.github.io/spectron/spectron) and spoken QA dataset (https://github.com/google-research-datasets/LLAMA1-Test-Set).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4d72b33-e666-41a5-b633-13b7474de3f8Cited by top-tier papers31
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationWenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen et al.NeurIPS 2025 · 43 citations
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language ModelsCheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin et al.ICLR 2026 · 38 citations
- Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI FeedbackGuan-Ting Lin, Prashanth Gurunath Shivakumar, Aditya Gourav, Yile Gu et al.ACL 2025 · 33 citations
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech InteractionShu-Wen Yang, Ming Tu, Ting-Wei Liu, Xinghua Qu et al.ICLR 2026 · 29 citations
- Paralinguistics-Aware Speech-Empowered Large Language Models for Natural ConversationHeeseung Kim, Soonshin Seo, Kyeongseok Jeong, Ohsung Kwon et al.NeurIPS 2024 · 28 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
Related papers
- Towards True Speech-to-Speech Models Without Text GuidanceXingjian Zhao, Zhe Xu, Luozhijie Jin, Yang Wang et al.ICLR 2026 · 8 citations
- Scaling Speech-Text Pre-training with Synthetic Interleaved DataAohan Zeng, Zhengxiao Du, Mingdao Liu, Lei Zhang et al.ICLR 2025
- Introducing Semantics into Speech EncodersDerek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim et al.ACL 2023 · 2 citations
- Data-Centric Lessons To Improve Speech-Language PretrainingVishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang et al.ICLR 2026 · 3 citations
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language ModelsGuangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen et al.ICML 2024 · 92 citations
