Benchmarking Audio Visual Segmentation for Long-Untrimmed Videos
Chen Liu, Peike Patrick Li, Qingtao Yu, Hongwei Sheng, Dadong Wang, Lincheng Li, Xin Yu
摘要
Existing audio-visual segmentation datasets typically focus on short-trimmed videos with only one pixel-map annotation for a per-second video clip. In contrast for untrimmed videos the sound duration start- and end-sounding time positions and visual deformation of audible objects vary significantly. Therefore we observed that current AVS models trained on trimmed videos might struggle to segment sounding objects in long videos. To investigate the feasibility of grounding audible objects in videos along both temporal and spatial dimensions we introduce the Long-Untrimmed Audio-Visual Segmentation dataset (LU-AVS) which includes precise frame-level annotations of sounding emission times and provides exhaustive mask annotations for all frames. Considering that pixel-level annotations are difficult to achieve in some complex scenes we also provide the bounding boxes to indicate the sounding regions. Specifically LU-AVS contains 10M mask annotations across 6.6K videos and 11M bounding box annotations across 7K videos. Compared with the existing datasets LU-AVS videos are on average 4 8 times longer with the silent duration being 3 15 times greater. Furthermore we try our best to adapt some baseline models that were originally designed for audio-visual-relevant tasks to examine the challenges of our newly curated LU-AVS. Through comprehensive evaluation we demonstrate the challenges of LU-AVS compared to the ones containing trimmed videos. Therefore LU-AVS provides an ideal yet challenging platform for evaluating audio-visual segmentation and localization on untrimmed long videos. The dataset is publicly available at: https://yenanliu.github.io/LU-AVS/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SAM Audio: Segment Anything in AudioBowen Shi, Andros Tjandra, John Hoffman, Helin Wang 等ICML 2026 · 被引用 35 次
- Implicit Counterfactual Learning for Audio-Visual SegmentationMingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang 等ICCV 2025 · 被引用 3 次
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual ScenariosLu Zhu, Tiantian Geng, Yangye Chen, Teng Wang 等AAAI 2026 · 被引用 1 次
- Beyond Single-View Sufficiency: CVBench for Cross-View Human UnderstandingTianchen Guo, Chen Liu, Xin YuCVPR 2026
它引用的顶会 Paper28
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
相关 Paper
- Audio-Visual Instance SegmentationRuohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu 等CVPR 2025
- Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and BaselineTiantian Geng, Teng Wang, Jinming Duan, Runmin Cong 等CVPR 2023
- Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsChen Liu, Peike Patrick Li, Xingqun Qi, Hu Zhang 等ACM MM 2023 · 被引用 33 次
- Open-Vocabulary Audio-Visual Semantic SegmentationRuohao Guo, Liao Qu, Dantong Niu, Yanyu Qi 等ACM MM 2024 · 被引用 4 次
- LVOS: A Benchmark for Long-term Video Object SegmentationLingyi Hong, Wenchao Chen, Zhongying Liu, Wei Zhang 等ICCV 2023 · 被引用 89 次
