TellWhisper: Tell Whisper Who Speaks When
Yifan Hu, Peiji Yang, Zhisheng Wang, Yicheng Zhong, Rui Liu
摘要
Multi-speaker automatic speech recognition (MASR) aims to predict "who spoke when and what" from multi-speaker speech, a key technology for multi-party dialogue understanding. However, most existing approaches decouple temporal modeling and speaker modeling when addressing "when" and "who": some inject speaker cues before encoding (e.g., speaker masking), which can cause irreversible information loss; others fuse identity by mixing speaker posteriors after encoding, which may entangle acoustic content with speaker identity. This separation is brittle under rapid turntaking and overlapping speech, often leading to degraded performance. To address these limitations, we propose TellWhisper, a unified framework that jointly models speaker identity and temporal within the speech encoder. Specifically, we design TS-RoPE, a time-speaker rotary positional encoding: time coordinates are derived from frame indices, while speaker coordinates are derived from speaker activity and pause cues. By applying region-specific rotation angles, the model explicitly captures per-speaker continuity, speaker-turn transitions, and state dynamics, enabling the attention mechanism to simultaneously attend to "when" and "who". Moreover, to estimate framelevel speaker activity, we develop Hyper-SD, which casts speaker classification in hyperbolic space to enhance inter-class separation and refine speaker-activity estimates. Extensive experiments demonstrate the effectiveness of the proposed approach. The project webpage is available at https://walker-hyf.github. io/TellWhisper .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Hyperbolic Neural Networks++Ryohei Shimizu, Yusuke Mukuta, Tatsuya HaradaICLR 2021 · 被引用 791 次
- Zipformer: A faster and better encoder for automatic speech recognitionZengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang 等ICLR 2024 · 被引用 155 次
- MoonCast: High-Quality Zero-Shot Podcast GenerationZeqian Ju, Dongchao Yang, Kai Shen, Yichong Leng 等NeurIPS 2025 · 被引用 32 次
- SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsHan Yin, Yafeng Chen, Chong Deng, Luyao Cheng 等AAAI 2026 · 被引用 18 次
- Hyperbolic Space with Hierarchical Margin Boosts Fine-Grained Learning from Coarse LabelsShu-Lin Xu, Yifan Sun, Faen Zhang, Anqi Xu 等NeurIPS 2023 · 被引用 16 次
相关 Paper
- WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text AdaptabilityYongkang Yin, Yuexian ZouAAAI 2026
- TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal GroundingMingyue Huo, Yiwen Shao, Yuheng ZhangACL 2026 · 被引用 12 次
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi 等CVPR 2020
- Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn DialogueLongxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou 等AAAI 2021 · 被引用 57 次
- Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field RecordingsChenyu Yang, Mengxi Chen, Yanfeng Wang, Yu WangACM MM 2023 · 被引用 2 次
