Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross- Modal Denoising Networks
Wenwen Pan, Haonan Shi, Zhou Zhao, Jieming Zhu, Xiuqiang He, Zhigeng Pan, Lianli Gao, Jun Yu, Fei Wu, Qi Tian
Abstract
Audio-Guided video object segmentation is a challenging problem in visual analysis and editing, which automatically separates foreground objects from the background in a video sequence according to the referring audio expressions. However, existing referring video object segmentation works mainly focus on the guidance of text-based referring expressions, due to the lack of modeling the semantic representations of audio-video interaction contents. In this paper, we consider the problem of audio-guided video semantic segmentation from the viewpoint of end-to-end denoising encoder-decoder network learning. We propose the wavelet-based encoder network to learn the cross-modal representations of the video contents with audio-form queries. Specifically, we adopt the multi-head cross-modal attention layers to explore the potential relations of video and query contents. A 2-dimension discrete wavelet trans-form is merged into the transformer encoder to decompose the audio-video features. Next, we maximize mutual information between the encoded features and multi-modal features after cross-modal attention layers to enhance the au-dio guidance. Then, a self attention-free decoder network is developed to generate the target masks with frequency-domain transforms. In addition, we construct the first large-scale audio-guided video semantic segmentation dataset. The extensive experiments show the effectiveness of our method <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Code is available at: https://github.com/asudahkzj/Wnet.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a6b18e9-9e84-4d49-81a0-e089f3e31840Cited by top-tier papers6
- Learning Temporal-Ordered Representation for Spike Streams Based on Discrete Wavelet TransformsJiyuan Zhang, Shanshan Jia, Zhaofei Yu, Tiejun HuangAAAI 2023 · 33 citations
- Curriculum-Listener: Consistency- and Complementarity-Aware Audio-Enhanced Temporal Sentence GroundingHoulun Chen, Xin Wang, Xiaohan Lan, Hong Chen et al.ACM MM 2023 · 13 citations
- Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and TextXiang Li, Jinglu Wang, Xiaohao Xu, Muqiao Yang et al.EMNLP 2023 · 6 citations
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual SegmentationKaining Ying, Henghui Ding, Guangquan Jie, Yu-Gang JiangICCV 2025 · 3 citations
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment RetrievalJunan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu et al.ACM MM 2025 · 3 citations
Builds on12
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler DivergenceXue Yang, Xiaojiang Yang, Jirui Yang, Qi Ming et al.NeurIPS 2021 · 603 citations
- Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language QueryHao Wang, Cheng Deng, Junchi Yan, Dacheng TaoICCV 2019 · 89 citations
Related papers
- Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationShilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen et al.AAAI 2024 · 67 citations
- Multi-Attention Network for Compressed Video Referring Object SegmentationWeidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han et al.ACM MM 2022 · 49 citations
- AVSegFormer: Audio-Visual Segmentation with TransformerShengyi Gao, Zhe Chen, Guo Chen, Wenhai Wang et al.AAAI 2024 · 96 citations
- TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual SegmentationAbduljalil Radman, Jorma LaaksonenCVPR 2025
- Cross-Modal Relation-Aware Networks for Audio-Visual Event LocalizationHaoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan et al.ACM MM 2020 · 97 citations
