SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li
Abstract
The Speaker Diarization and Recognition (SDR) task aims to predict ``who spoke when and what'' within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3b2ac7f-e371-40b7-bcbc-e29a78097ef9Cited by top-tier papers4
- TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal GroundingMingyue Huo, Yiwen Shao, Yuheng ZhangACL 2026 · 12 citations
- VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language ModelsYuxiang Wang, HongYu Liu, Dekun Chen, Xueyao Zhang et al.ICLR 2026 · 5 citations
- TellWhisper: Tell Whisper Who Speaks WhenYifan Hu, Peiji Yang, Zhisheng Wang, Yicheng Zhong et al.ACL 2026 · 1 citation
- Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party InterruptionsDongwook Lee, Eunwoo Song, Che Hyun Lee, Heeseung Kim et al.ACL 2026
Builds on1
Related papers
- WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text AdaptabilityYongkang Yin, Yuexian ZouAAAI 2026
- DNCASR: End-to-End Training for Speaker-Attributed ASRXianrui Zheng, Chao Zhang, Philip C. WoodlandACL 2025 · 6 citations
- CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker DiarizationLiangbin Huang, Xiaohua Liao, Chaoqun Cui, Shijing Wang et al.CVPR 2026
- Reasoning LLM Improves Speaker Recognition in Long-form TV DramasYuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu et al.ICML 2026
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 74 citations
