WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability
Yongkang Yin, Yuexian Zou
摘要
Speaker diarization is a fundamental task in speech processing aims to determine 'who speaks when'. When combined with ASR, it enables speaker-labeled transcription with broad practical value. Most existing methods rely on frame-level classification, but the high cost of annotating mixed-speaker audio limits the availability of large-scale, accurately labeled datasets. As a result, even state-of-the-art models struggle with imprecise speaker boundary detection and semantic segmentation errors, which degrade timestamp accuracy and downstream ASR performance. To address these challenges, we propose WhisperDiari, a unified framework for speaker diarization and ASR. We first construct LibriDiari, a dataset derived from LibriSpeech, containing 2–4 speaker mixed audio annotated with transcripts and speaker labels. WhisperDiari builds on the Whisper model, incorporating speaker adapters and Speaker Similarity Matrix Supervision to enhance speaker representation. In addition, a dedicated speaker decoder fuses speaker embeddings with contextual semantics from Whisper's decoder, enabling token-level diarization. This design effectively resolves segmentation ambiguity, aligns diarization with semantic units, and jointly models 'who speaks what and when', producing accurate, timestamped transcripts. We train the model on LibriDiari and evaluate it on both LibriDiari and the real-world AMI corpus. Experimental results demonstrate that WhisperDiari consistently outperforms state-of-the-art open-source baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Recent Advances in Speech Language Models: A SurveyWenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng 等ACL 2025
- Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text SystemsTaejin Park, Ivan Medennikov, Kunal Dhawan, Weiqing Wang 等ICML 2025
相关 Paper
- TellWhisper: Tell Whisper Who Speaks WhenYifan Hu, Peiji Yang, Zhisheng Wang, Yicheng Zhong 等ACL 2026 · 被引用 1 次
- Whisper-UT: A Unified Translation Framework for Speech and TextCihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz 等EMNLP 2025
- SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsHan Yin, Yafeng Chen, Chong Deng, Luyao Cheng 等AAAI 2026 · 被引用 18 次
- TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal GroundingMingyue Huo, Yiwen Shao, Yuheng ZhangACL 2026 · 被引用 12 次
- Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field RecordingsChenyu Yang, Mengxi Chen, Yanfeng Wang, Yu WangACM MM 2023 · 被引用 2 次
