The Right to Talk: An Audio-Visual Transformer Approach
Thanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham, Bhiksha Raj, Ngan Le, Khoa Luu
摘要
Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker's utterances) remains a challenging task. Although some prior methods have partially addressed this task, there still remain some limitations. Firstly, a direct association of Audio and Visual features may limit the correlations to be extracted due to different modalities. Secondly, the relationship across temporal segments helping to maintain the consistency of localization, separation and conversation contexts is not effectively exploited. Finally, the interactions between speakers that usually contain the tracking and anticipatory decisions about transition to a new speaker is usually ignored. Therefore, this work introduces a new Audio-Visual Transformer approach to the problem of localization and highlighting the main speaker in both audio and visual channels of a multi-speaker conversation video in the wild. The proposed method exploits different types of correlations presented in both visual and audio signals. The temporal audio-visual relationships across spatial-temporal space are anticipated and optimized via the self-attention mechanism in a Transformer structure. Moreover, a newly collected dataset is introduced for the main speaker detection. To the best of our knowledge, it is one of the first studies that is able to automatically localize and highlight the main speaker in both visual and audio channels in multi-speaker conversation videos. Table 1. Comparisons of our proposed approach and other modeling methods. Sound Source Localization (SSL)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Visual Acoustic MatchingChangan Chen, Ruohan Gao, Paul Calamia, Kristen GraumanCVPR 2022 · 被引用 42 次
- Egocentric Deep Multi-Channel Audio-Visual Active Speaker LocalizationHao Jiang, Calvin Murdock, Vamsi Krishna IthapuCVPR 2022 · 被引用 39 次
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationSamuel Pegg, Kai Li, Xiaolin HuICLR 2024 · 被引用 13 次
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion LearningThanh-Dat Truong, Christophe Bobda, Nitin Agarwal, Khoa LuuNeurIPS 2025 · 被引用 6 次
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 被引用 5 次
它引用的顶会 Paper7
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Voice Separation with an Unknown Number of Multiple SpeakersEliya Nachmani, Yossi Adi, Lior WolfICML 2020 · 被引用 186 次
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan 等NeurIPS 2020 · 被引用 156 次
- DyGLIP: A Dynamic Graph Model With Link Prediction for Accurate Multi-Camera Multiple Object TrackingKha Gia Quach, Pha A. Nguyen, Huu Le, Thanh-Dat Truong 等CVPR 2021
相关 Paper
- AV-Dialog: Spoken Dialogue Models with Audio-Visual InputTuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath GollakotaACL 2026 · 被引用 1 次
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 等ACM MM 2021 · 被引用 154 次
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi 等CVPR 2020
- CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video SegmentationKexin Li, Zongxin Yang, Lei Chen, Yi Yang 等ACM MM 2023 · 被引用 58 次
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric PerspectiveWenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla 等CVPR 2024
