TellWhisper: Tell Whisper Who Speaks When
Yifan Hu, Peiji Yang, Zhisheng Wang, Yicheng Zhong, Rui Liu
Abstract
Multi-speaker automatic speech recognition (MASR) aims to predict "who spoke when and what" from multi-speaker speech, a key technology for multi-party dialogue understanding. However, most existing approaches decouple temporal modeling and speaker modeling when addressing "when" and "who": some inject speaker cues before encoding (e.g., speaker masking), which can cause irreversible information loss; others fuse identity by mixing speaker posteriors after encoding, which may entangle acoustic content with speaker identity. This separation is brittle under rapid turntaking and overlapping speech, often leading to degraded performance. To address these limitations, we propose TellWhisper, a unified framework that jointly models speaker identity and temporal within the speech encoder. Specifically, we design TS-RoPE, a time-speaker rotary positional encoding: time coordinates are derived from frame indices, while speaker coordinates are derived from speaker activity and pause cues. By applying region-specific rotation angles, the model explicitly captures per-speaker continuity, speaker-turn transitions, and state dynamics, enabling the attention mechanism to simultaneously attend to "when" and "who". Moreover, to estimate framelevel speaker activity, we develop Hyper-SD, which casts speaker classification in hyperbolic space to enhance inter-class separation and refine speaker-activity estimates. Extensive experiments demonstrate the effectiveness of the proposed approach. The project webpage is available at https://walker-hyf.github. io/TellWhisper .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc14d230-0251-4e2b-82c8-cc5466d34a25Builds on9
- Hyperbolic Neural Networks++Ryohei Shimizu, Yusuke Mukuta, Tatsuya HaradaICLR 2021 · 791 citations
- Zipformer: A faster and better encoder for automatic speech recognitionZengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang et al.ICLR 2024 · 155 citations
- MoonCast: High-Quality Zero-Shot Podcast GenerationZeqian Ju, Dongchao Yang, Kai Shen, Yichong Leng et al.NeurIPS 2025 · 32 citations
- SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsHan Yin, Yafeng Chen, Chong Deng, Luyao Cheng et al.AAAI 2026 · 18 citations
- Hyperbolic Space with Hierarchical Margin Boosts Fine-Grained Learning from Coarse LabelsShu-Lin Xu, Yifan Sun, Faen Zhang, Anqi Xu et al.NeurIPS 2023 · 16 citations
Related papers
- WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text AdaptabilityYongkang Yin, Yuexian ZouAAAI 2026
- TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal GroundingMingyue Huo, Yiwen Shao, Yuheng ZhangACL 2026 · 12 citations
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi et al.CVPR 2020
- Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn DialogueLongxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou et al.AAAI 2021 · 57 citations
- Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field RecordingsChenyu Yang, Mengxi Chen, Yanfeng Wang, Yu WangACM MM 2023 · 2 citations
