WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability
Yongkang Yin, Yuexian Zou
Abstract
Speaker diarization is a fundamental task in speech processing aims to determine 'who speaks when'. When combined with ASR, it enables speaker-labeled transcription with broad practical value. Most existing methods rely on frame-level classification, but the high cost of annotating mixed-speaker audio limits the availability of large-scale, accurately labeled datasets. As a result, even state-of-the-art models struggle with imprecise speaker boundary detection and semantic segmentation errors, which degrade timestamp accuracy and downstream ASR performance. To address these challenges, we propose WhisperDiari, a unified framework for speaker diarization and ASR. We first construct LibriDiari, a dataset derived from LibriSpeech, containing 2–4 speaker mixed audio annotated with transcripts and speaker labels. WhisperDiari builds on the Whisper model, incorporating speaker adapters and Speaker Similarity Matrix Supervision to enhance speaker representation. In addition, a dedicated speaker decoder fuses speaker embeddings with contextual semantics from Whisper's decoder, enabling token-level diarization. This design effectively resolves segmentation ambiguity, aligns diarization with semantic units, and jointly models 'who speaks what and when', producing accurate, timestamped transcripts. We train the model on LibriDiari and evaluate it on both LibriDiari and the real-world AMI corpus. Experimental results demonstrate that WhisperDiari consistently outperforms state-of-the-art open-source baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0e448a9-34a1-486d-87e4-7595e831bef5Builds on3
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Recent Advances in Speech Language Models: A SurveyWenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng et al.ACL 2025
- Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text SystemsTaejin Park, Ivan Medennikov, Kunal Dhawan, Weiqing Wang et al.ICML 2025
Related papers
- TellWhisper: Tell Whisper Who Speaks WhenYifan Hu, Peiji Yang, Zhisheng Wang, Yicheng Zhong et al.ACL 2026 · 1 citation
- Whisper-UT: A Unified Translation Framework for Speech and TextCihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz et al.EMNLP 2025
- SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsHan Yin, Yafeng Chen, Chong Deng, Luyao Cheng et al.AAAI 2026 · 18 citations
- TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal GroundingMingyue Huo, Yiwen Shao, Yuheng ZhangACL 2026 · 12 citations
- Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field RecordingsChenyu Yang, Mengxi Chen, Yanfeng Wang, Yu WangACM MM 2023 · 2 citations
