Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Henghui Du, Guangyao Li, Chang Zhou, Chunjie Zhang, Alan Zhao, Di Hu
摘要
Please recognize the category of object that makes the sound and then output the location Spatial localization coordinates. Please describe the events and time range that occurred in the video. A red car appears from a distance and drives down a dirt road, kicking up dust and creating a cloud of smoke. The car is visible and audible from the 4th to the 9th second. Temporal localization Please segment out the object that makes sound on the left. Please segment out the sounding object. Pixel-level understanding In the video, three people are playing musical instruments in front of a Christmas tree. The man on the left is playing the cello, the man in the middle is playing the violin, and the man on the right is playing the piano. At the 2nd second, the piano sounds first. Then, starting from the 4th second, three instruments play together. The instrument on the left of the piano is the cello. So the answer is cello. Spatio-temporal reasoning What is the left instrument of the first sounding instrument? Audio-visual scene understanding task AVE/AVVP/AVQA/AVS/ARIG/… MLLMs Temporal localization Spatial localization Spatio-temporal reasoning Pixel-level understanding 20241114 Figure 1. We present Crab, a unified audio-visual scene understanding model with explicit cooperation, which can complete various audio-visual tasks. It is trained on an instruction-tuning dataset with explicit reasoning process, which clarifies the cooperative relationship among tasks. Furthermore, to alleviate the interference caused by the learning process of complex audiovisual data and facilitate concrete cooperation, an interaction-aware LoRA structure is designed to enable the model focus on different aspects of data interaction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsLidong Lu, Guo Chen, Zhu Wei, Zhiqi Li 等CVPR 2026 · 被引用 23 次
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video HaystacksSanjoy Chowdhury, Mohamed Elmoghany, Yohan Abeysinghe, Junjie Fei 等NeurIPS 2025 · 被引用 14 次
- Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationJinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang 等AAAI 2026 · 被引用 5 次
- MoXaRt: Audio-Visual Object-Guided Sound Interaction for XRTianyu Xu, Sieun Kim, Qianhui Zheng, Ruoyu Xu 等CHI 2026 · 被引用 1 次
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual ScenariosLu Zhu, Tiantian Geng, Yangye Chen, Teng Wang 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu 等ICML 2023 · 被引用 568 次
相关 Paper
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesXingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer 等ICML 2025
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language ModelsHewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang 等CVPR 2026 · 被引用 3 次
- Consistent and Controllable Image Animation with Motion Diffusion ModelsXin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen 等CVPR 2025
- CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus AttentionsHaicheng Liao, Haoyu Sun, Huanming Shen, Chengyue Wang 等ACM MM 2024 · 被引用 10 次
