CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video Segmentation
Kexin Li, Zongxin Yang, Lei Chen, Yi Yang, Jun Xiao
Abstract
Audio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects within image frames and ensure the maps faithfully adheres to the given audio, such as identifying and segmenting a singing person in a video. However, existing methods exhibit two limitations: 1) they address video temporal features and audio-visual interactive features separately, disregarding the inherent spatial-temporal dependence of combined audio and video, and 2) they inadequately introduce audio constraints and object-level information during the decoding stage, resulting in segmentation outcomes that fail to comply with audio directives. To tackle these issues, we propose a decoupled audio-video transformer that combines audio and video features from their respective temporal and spatial dimensions, capturing their combined dependence. To optimize memory consumption, we design a block, which, when stacked, enables capturing audio-visual fine-grained combinatorial-dependence in a memory-efficient manner. Additionally, we introduce audio-constrained queries during the decoding phase. These queries contain rich object-level information, ensuring the decoded mask adheres to the sounds. Experimental results confirm our approach's effectiveness, with our framework achieving a new SOTA performance on all three datasets using two backbones. The code is available at https://github.com/aspirinone/CATR.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4984f2d1-eda1-42d4-8cc4-be5f9b97dc37Cited by top-tier papers31
- Efficient Emotional Adaptation for Audio-Driven Talking-Head GenerationYuan Gan, Zongxin Yang, Xihang Yue, Lingyun Sun et al.ICCV 2023 · 111 citations
- Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object SegmentationShaofei Huang, Rui Ling, Hongyu Li, Tianrui Hui et al.AAAI 2025 · 24 citations
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation LearningLuoyi Sun, Xuenan Xu, Mengyue Wu, Weidi XieACM MM 2024 · 23 citations
- JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh RecoveryJiahao Li, Zongxin Yang, Xiaohan Wang, Jianxin Ma et al.ICCV 2023 · 22 citations
- Boosting Audio Visual Question Answering via Key Semantic-Aware CuesGuangyao Li, Henghui Du, Di HuACM MM 2024 · 16 citations
Builds on27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Revisiting Audio-Visual Segmentation with Vision-Centric TransformerShaofei Huang, Rui Ling, Tianrui Hui, Hongyu Li et al.CVPR 2025
- AVSegFormer: Audio-Visual Segmentation with TransformerShengyi Gao, Zhe Chen, Guo Chen, Wenhai Wang et al.AAAI 2024 · 96 citations
- QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic DecompositionXiang Li, Jinglu Wang, Xiaohao Xu, Xiulian Peng et al.CVPR 2024
- Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual SegmentationQi Yang, Xing Nie, Tong Li, Pengfei Gao et al.CVPR 2024 · 9 citations
- SelM: Selective Mechanism based Audio-Visual SegmentationJiaxu Li, Songsong Yu, Yifan Wang, Lijun Wang et al.ACM MM 2024 · 5 citations
