Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Henghui Du, Guangyao Li, Chang Zhou, Chunjie Zhang, Alan Zhao, Di Hu
Abstract
Please recognize the category of object that makes the sound and then output the location Spatial localization coordinates. Please describe the events and time range that occurred in the video. A red car appears from a distance and drives down a dirt road, kicking up dust and creating a cloud of smoke. The car is visible and audible from the 4th to the 9th second. Temporal localization Please segment out the object that makes sound on the left. Please segment out the sounding object. Pixel-level understanding In the video, three people are playing musical instruments in front of a Christmas tree. The man on the left is playing the cello, the man in the middle is playing the violin, and the man on the right is playing the piano. At the 2nd second, the piano sounds first. Then, starting from the 4th second, three instruments play together. The instrument on the left of the piano is the cello. So the answer is cello. Spatio-temporal reasoning What is the left instrument of the first sounding instrument? Audio-visual scene understanding task AVE/AVVP/AVQA/AVS/ARIG/… MLLMs Temporal localization Spatial localization Spatio-temporal reasoning Pixel-level understanding 20241114 Figure 1. We present Crab, a unified audio-visual scene understanding model with explicit cooperation, which can complete various audio-visual tasks. It is trained on an instruction-tuning dataset with explicit reasoning process, which clarifies the cooperative relationship among tasks. Furthermore, to alleviate the interference caused by the learning process of complex audiovisual data and facilitate concrete cooperation, an interaction-aware LoRA structure is designed to enable the model focus on different aspects of data interaction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6bfc693-a2d1-4830-8cf1-654372f34753Cited by top-tier papers9
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsLidong Lu, Guo Chen, Zhu Wei, Zhiqi Li et al.CVPR 2026 · 23 citations
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video HaystacksSanjoy Chowdhury, Mohamed Elmoghany, Yohan Abeysinghe, Junjie Fei et al.NeurIPS 2025 · 14 citations
- Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationJinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang et al.AAAI 2026 · 5 citations
- MoXaRt: Audio-Visual Object-Guided Sound Interaction for XRTianyu Xu, Sieun Kim, Qianhui Zheng, Ruoyu Xu et al.CHI 2026 · 1 citation
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual ScenariosLu Zhu, Tiantian Geng, Yangye Chen, Teng Wang et al.AAAI 2026 · 1 citation
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu et al.ICML 2023 · 568 citations
Related papers
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesXingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer et al.ICML 2025
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language ModelsHewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang et al.CVPR 2026 · 3 citations
- Consistent and Controllable Image Animation with Motion Diffusion ModelsXin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen et al.CVPR 2025
- CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus AttentionsHaicheng Liao, Haoyu Sun, Huanming Shen, Chengyue Wang et al.ACM MM 2024 · 10 citations
