Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field Recordings
Chenyu Yang, Mengxi Chen, Yanfeng Wang, Yu Wang
摘要
Audio-visual speaker diarization refers to the task of identifying "who spoke when" by using both audio and video data. Although previous fusion-based approaches have shown exceptional performance over audio-only methods, they have mainly focused on high-quality data and have not accounted for the impacts of acoustic noise or missing faces. To address these limitations, we propose a novel uncertainty-aware end-to-end audio-visual speaker diarization (UAV-SD) approach in this paper. Our approach leverages both framewise inter- and intra-modal confidence to achieve more effective and robust speaker diarization. By taking into account the uncertainty of the data, UAV-SD can achieve better diarization performance even in noisy or low-quality recordings. Additionally, our approach is compatible with multi-channel audio signals without the need to retrain the model, making it a more versatile solution. To evaluate the effectiveness of our approach, we conduct extensive experiments on the Multi-modal Information Based Speech Processing (MISP) 2022 Challenge datasets which consist of far-field audio and video data. The results show that UAV-SD is able to yield significant performance gains compared to baseline methods for both single and multi-channel data, demonstrating its effectiveness in real-world scenarios.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng 等ACM MM 2022 · 被引用 34 次
- WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text AdaptabilityYongkang Yin, Yuexian ZouAAAI 2026
- Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party ConversationLuyao Cheng, Hui Wang, Chong Deng, Siqi Zheng 等ACL 2025
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech RepresentationQiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu 等AAAI 2024 · 被引用 17 次
