Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation
Xinmeng Xu, Haoran Xie, Xiaohui Tao, Lin Li, S. Joe Qin
摘要
Current audio-visual speech separation (AVSS) models typically rely on implicit multimodal fusion, but the absence of explicit modality alignment and reliability modeling often causes semantic misalignment and contaminates speech representations. The brain addresses this with a hierarchy: top-down auditory selection uses visual priors to maintain target-consistent acoustics, while bottom-up cross-modal compensation integrates temporally aligned articulatory cues to reconstruct and stabilize speech. Guided by this principle, we present Neuro-SCNet, an AVSS architecture that makes selection and compensation explicit and reliability-aware. The Auditory Selection Mechanism applies top-down, visually guided gain along the audio pathway to isolate target time-frequency units and suppress distractors. The module preserves the auditory trace with an identity bypass and adds controlled visual refinements via a residual path. A synchrony-driven gate reduces the influence of low-confidence visual cues. Additionally, a lightweight pre-alignment for visual feature pre-processing estimates and corrects small temporal offsets, and a compact magnitude-phase encoder is used to preserve fine acoustic detail to stabilize reconstruction. Evaluations on LRS2, LRS3, and VoxCeleb2 show state-of-the-art separation with improved efficiency, supporting the value of explicit selection and reliability-aware compensation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech SeparationUi-Hyeop Shin, Sangyoun Lee, Taehan Kim, Hyung-Min ParkNeurIPS 2024 · 被引用 46 次
- IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationKai Li, Runxuan Yang, Fuchun Sun, Xiaolin HuICML 2024 · 被引用 28 次
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationSamuel Pegg, Kai Li, Xiaolin HuICLR 2024 · 被引用 13 次
- RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual CuesTianrui Pan, Jie Liu, Bohan Wang, Jie Tang 等ACM MM 2024 · 被引用 3 次
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
相关 Paper
- SelM: Selective Mechanism based Audio-Visual SegmentationJiaxu Li, Songsong Yu, Yifan Wang, Lijun Wang 等ACM MM 2024 · 被引用 5 次
- Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringJoanna Hong, Minsu Kim, Jeongsoo Choi, Yong Man RoCVPR 2023
- CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveJunwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang 等CVPR 2023
- Reading to Listen at the Cocktail Party: Multi-Modal Speech SeparationAkam Rahimi, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 21 次
- Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual ScenesHyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda SenocakCVPR 2025
