SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li
摘要
The Speaker Diarization and Recognition (SDR) task aims to predict ``who spoke when and what'' within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal GroundingMingyue Huo, Yiwen Shao, Yuheng ZhangACL 2026 · 被引用 12 次
- VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language ModelsYuxiang Wang, HongYu Liu, Dekun Chen, Xueyao Zhang 等ICLR 2026 · 被引用 5 次
- TellWhisper: Tell Whisper Who Speaks WhenYifan Hu, Peiji Yang, Zhisheng Wang, Yicheng Zhong 等ACL 2026 · 被引用 1 次
- Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party InterruptionsDongwook Lee, Eunwoo Song, Che Hyun Lee, Heeseung Kim 等ACL 2026
它引用的顶会 Paper1
相关 Paper
- WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text AdaptabilityYongkang Yin, Yuexian ZouAAAI 2026
- DNCASR: End-to-End Training for Speaker-Attributed ASRXianrui Zheng, Chao Zhang, Philip C. WoodlandACL 2025 · 被引用 6 次
- CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker DiarizationLiangbin Huang, Xiaohua Liao, Chaoqun Cui, Shijing Wang 等CVPR 2026
- Reasoning LLM Improves Speaker Recognition in Long-form TV DramasYuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu 等ICML 2026
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 被引用 74 次
