Pano-AVQA: Grounded Audio-Visual Question Answering on 360° Videos
Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, Gunhee Kim
摘要
360° videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond predetermined normal field of views and displays distinctive spatial relations on a sphere. However, previous benchmark tasks for panoramic videos are still limited to evaluate the semantic understanding of audio-visual relationships or spherical spatial property in surroundings. We propose a novel benchmark named Pano-AVQA as a large-scale grounded audio-visual question answering dataset on panoramic videos. Using 5.4K 360° video clips harvested online, we collect two types of novel question-answer pairs with bounding-box grounding: spherical spatial relation QAs and audio-visual relation QAs. We train several transformer-based models from Pano-AVQA, where the results suggest that our proposed spherical spatial embeddings and multimodal training objectives fairly contribute to a better semantic understanding of the panoramic surroundings on the dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question AnsweringJie Ma, Min Hu, Pinghui Wang, Wangchun Sun 等NeurIPS 2024 · 被引用 31 次
- Object-Aware Adaptive-Positivity Learning for Audio-Visual Question AnsweringZhangbin Li, Dan Guo, Jinxing Zhou, Jing Zhang 等AAAI 2024 · 被引用 30 次
- Patch-level Sounding Object Tracking for Audio-Visual Question AnsweringZhangbin Li, Jinxing Zhou, Jing Zhang, Shengeng Tang 等AAAI 2025 · 被引用 20 次
- Audio-Visual Adaptive Fusion Network for Question Answering Based on Contrastive LearningXujian Zhao, Yixin Wang, Peiquan JinAAAI 2025 · 被引用 4 次
- Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream TasksColin Leong, Joshua Nemecek, Jacob Mansdorfer, Anna Filighera 等EMNLP 2022 · 被引用 3 次
它引用的顶会 Paper8
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- WoodScape: A Multi-Task, Multi-Camera Fisheye Dataset for Autonomous DrivingSenthil Kumar Yogamani, Christian Witt, Hazem Rashed, Sanjaya Nayak 等ICCV 2019 · 被引用 325 次
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
相关 Paper
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 等CVPR 2022 · 被引用 101 次
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 60 次
- PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement LearningZekai Lin, Xu ZhengCVPR 2026 · 被引用 7 次
- More than the Sum: Panorama-Language Models for Adverse Omni-ScenesWeijia Fan, Ruiping Liu, Jiale Wei, Yufan Chen 等CVPR 2026 · 被引用 5 次
- ScanQA: 3D Question Answering for Spatial Scene UnderstandingDaichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki KawanabeCVPR 2022 · 被引用 135 次
