Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic Environments
Difei Gao, Ruiping Wang, Ziyi Bai, Xilin Chen
摘要
Visual understanding goes well beyond the study of images or videos on the web. To achieve complex tasks in volatile situations, the human can deeply understand the environment, quickly perceive events happening around, and continuously track objects’ state changes, which are still challenging for current AI systems. To equip AI system with the ability to understand dynamic ENVironments, we build a video Question Answering dataset named Env-QA. Env-QA contains 23K egocentric videos, where each video is composed of a series of events about exploring and interacting in the environment. It also provides 85K questions to evaluate the ability of understanding the composition, layout, and state changes of the environment presented by the events in videos. Moreover, we propose a video QA model, Temporal Segmentation and Event Attention network (TSEA), which introduces event-level video representation and corresponding attention mechanisms to better extract environment information and answer questions. Comprehensive experiments demonstrate the effectiveness of our framework and show the formidable challenges of Env-QA in terms of long-term state tracking, multi-event temporal reasoning and event counting, etc.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li 等EMNLP 2022 · 被引用 70 次
- MoReVQA: Exploring Modular Reasoning Models for Video Question AnsweringJuhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho 等CVPR 2024 · 被引用 27 次
- Glance and Focus: Memory Prompting for Multi-Event Video Question AnsweringZiyi Bai, Ruiping Wang, Xilin ChenNeurIPS 2023 · 被引用 19 次
- GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented CollaborationsMuhammet Furkan Ilaslan, Chenan Song, Joya Chen, Difei Gao 等EMNLP 2023 · 被引用 6 次
- Encoding and Controlling Global Semantics for Long-form Video Question AnsweringThong Nguyen, Zhiyuan Hu, Xiaobao Wu, Cong-Duy Nguyen 等EMNLP 2024 · 被引用 3 次
它引用的顶会 Paper8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 被引用 183 次
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 被引用 173 次
相关 Paper
- Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQAHyounghun Kim, Zineng Tang, Mohit BansalACL 2020 · 被引用 31 次
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 被引用 47 次
- Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene UnderstandingYue Fan, Xiaojian Ma, Rongpeng Su, Jun Guo 等ICCV 2025 · 被引用 2 次
- AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityRengang Li, Cong Xu, Zhenhua Guo, Baoyu Fan 等ACM MM 2022 · 被引用 7 次
- Dynamic Spatio-Temporal Modular Network for Video Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 14 次
