Closing the Modality Reasoning Gap for Speech Large Language Models
Chaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu, Yan Lu, Jinyu Li, Zhizheng Wu
摘要
Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on speech inputs is markedly weaker than on text. This gap could be associated with representational drift across Transformer layers and behavior deviations in long-chain reasoning. To address this issue, we introduce TARS, a reinforcement-learning framework that aligns text-conditioned and speech-conditioned trajectories through an asymmetric reward design. The framework employs two dense and complementary signals: representation alignment, which measures layer-wise hidden-state similarity between speech- and text-conditioned trajectories, and behavior alignment, which evaluates semantic consistency between generated outputs and reference text completions. Experiments on challenging reasoning benchmarks, including MMSU and OBQA, show that our approach significantly narrows the modality reasoning gap and achieves state-of-the-art performance among 7B-scale Speech LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 等ICLR 2024 · 被引用 557 次
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky 等ICLR 2024 · 被引用 247 次
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang 等ICLR 2026 · 被引用 143 次
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 被引用 95 次
相关 Paper
- Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language ModelsBajian Xiang, Shuaijiang Zhao, Tingwei Guo, Wei ZouEMNLP 2025 · 被引用 6 次
- One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query RefinementYixiao Zhou, Dongzhou Cheng, Zhiliang Wu, Yi Yang 等ACL 2026 · 被引用 3 次
- Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement LearningShu Wu, Chenxing Li, Wenfu Wang, Hao Zhang 等AAAI 2026 · 被引用 4 次
- MARS-Sep: Multimodal-Aligned Reinforced Sound SeparationZihan Zhang, Xize Cheng, Zhennan Jiang, Dongjie Fu 等ICLR 2026 · 被引用 2 次
- Closing the Gap Between Text and Speech Understanding in LLMsSantiago Cuervo, Skyler Seto, Maureen de Seyssel, Richard He Bai 等ICLR 2026 · 被引用 18 次
