Multi-Factor Adaptive Vision Selection for Egocentric Video Question Answering
Haoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song, Yaowei Wang, Liqiang Nie
Abstract
The challenge of interpreting the world from a human perspective in Artificial Intelligence (AI) is particularly evident in egocentric video question answering, which grapples with issues like small object recognition, noise suppression, and spatial-temporal reasoning. To address these challenges, we introduce the Multi-Factor Adaptive vision Selection (MFAS) framework. MFAS integrates a patch partition and merging module for enhanced small object recognition, a priorguided patch selection module for noise suppression and focused analysis, and a hierarchical aggregation network to aggregate visual semantics guided by questions. Extensive experiments on several public egocentric datasets have validated the effectiveness and generalization of our framework. Code and data are available in https:// github.com/Hyu-Zhang/EgoVideoQA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51e4ea31-d2ee-464a-a14b-779f0543c9d7Cited by top-tier papers15
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen et al.NeurIPS 2024 · 104 citations
- Spatial Understanding from Videos: Structured Prompts Meet Simulation DataHaoyu Zhang, Meng Liu, Zaijing Li, Haokun Wen et al.NeurIPS 2025 · 31 citations
- Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic ManipulationZaijing Li, Bing Hu, Rui Shao, Gongwei Chen et al.CVPR 2026 · 23 citations
- Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video UnderstandingHaoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi et al.AAAI 2026 · 17 citations
- Fair Deepfake Detectors Can GeneralizeHarry Cheng, Ming-Hui Liu, Yangyang Guo, Tianyi Wang et al.NeurIPS 2025 · 11 citations
Builds on32
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.ICCV 2021 · 345 citations
- Egocentric Video-Language PretrainingKevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray et al.NeurIPS 2022 · 306 citations
Related papers
- Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic EnvironmentsDifei Gao, Ruiping Wang, Ziyi Bai, Xilin ChenICCV 2021 · 36 citations
- SegEQA: Video Segmentation Based Visual Attention for Embodied Question AnsweringHaonan Luo, Guosheng Lin, Zichuan Liu, Fayao Liu et al.ICCV 2019 · 29 citations
- EgoHierMask: Hierarchical Semantic-Prior Guided Masked Autoencoder for Egocentric Action RecognitionJiang Shao, Xinbo Zhao, Xiaochun Zou, Xiaolin YeACM MM 2025
- MAMS: Model-Agnostic Module Selection Framework for Video CaptioningSangho Lee, Il Yong Chun, Hogun ParkAAAI 2025 · 1 citation
- Where is my Wallet? Modeling Object Proposal Sets for Egocentric Visual Query LocalizationMengmeng Xu, Yanghao Li, Cheng-Yang Fu, Bernard Ghanem et al.CVPR 2023
