AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow Matching
Xize Cheng, Chenyuhao Wen, Slytherin Wang, Yongqi Wang, Zehan Wang, Rongjie Huang, Tao Jin, Zhou Zhao
Abstract
Video Query Sound Separation (VQSS) aims to isolate target sounds conditioned on visual queries while suppressing off-screen interference-a task central to audiovisual understanding. However, existing methods often fail under conditions of homogeneous interference and overlapping soundtracks, due to limited temporal modeling and weak audiovisual alignment. We propose AlignSep, the first generative VQSS model based on flow matching, designed to address common issues such as spectral holes and incomplete separation. To better capture crossmodal correspondence, we introduce a series of temporal consistency mechanisms that guide the vector field estimator toward learning robust audiovisual alignment, enabling accurate and resilient separation in complex scenes. As a multiconditioned generation task, VQSS presents unique challenges that differ fundamentally from traditional flow matching setups. We provide an in-depth analysis of these differences and their implications for generative modeling. To systematically evaluate performance under realistic and difficult conditions, we further construct VGGSound-Hard, a challenging benchmark composed entirely of separation cases with homogeneous interference and strong reliance on temporal visual cues. Extensive experiments across multiple benchmarks demonstrate that AlignSep achieves state-of-the-art performance both quantitatively and perceptually, validating its practical value for real-world applications. More results and audio examples are available at: https://AlignSep.github.io . * Equal Contribution • We revisit the task of video-queried sound separation (VQSS) and provide a detailed analysis of its unique challenges, including homogeneous interference, overlapping soundtracks, and the need for precise audio-visual temporal alignment. • We propose AlignSep, a novel generative temporal-aligned VQSS framework based on conditional flow matching, designed to robustly model multi-conditioned generation by leveraging temporal visual cues and preserving cross-modal consistency. • We introduce VGGSound-Hard, a new benchmark specifically curated to evaluate temporal alignment under real-world homogeneous interference, consisting of co-occurring on-/off-screen same-category sound sources. • Extensive experiments on three benchmarks-MUSIC-Clean, VGGSound-Clean, and VGGSound-Hard-demonstrate that AlignSep achieves state-of-the-art performance in both quantitative metrics and human perceptual scores (e.g., MOS), validating its effectiveness in real-world audiovisual separation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebb34b8f-8586-4e42-b6cd-b5b86688c014Builds on15
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss et al.NeurIPS 2020 · 227 citations
- Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion ModelsSimian Luo, Chuanhao Yan, Chenxu Hu, Hang ZhaoNeurIPS 2023 · 192 citations
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen SoundsEfthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey et al.ICLR 2021 · 83 citations
Related papers
- Cinematic Audio Source Separation Using Visual CuesKang Zhang, Suyeon Lee, Arda Senocak, Joon Son ChungCVPR 2026 · 1 citation
- VAFlow: Video-to-Audio Generation with Cross-Modality Flow MatchingXihua Wang, Xin Cheng, Yuyue Wang, Ruihua Song et al.ICCV 2025 · 6 citations
- OmniSep: Unified Omni-Modality Sound Separation with Query-MixupXize Cheng, Siqi Zheng, Zehan Wang, Minghui Fang et al.ICLR 2025
- Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio GenerationKang Zhang, Trung X. Pham, Suyeon Lee, Axi Niu et al.NeurIPS 2025 · 1 citation
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
