Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
Insu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh, Kyuhong Shim, Byonghyo Shim
Abstract
Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view offers fine-grained cues about user attention and hand-object interactions, its narrow field of view and lack of global context often lead to failures on spatially or contextually demanding queries. To address this, we introduce a framework that augments egocentric inputs with third-person (exocentric) views, providing complementary information such as global scene layout and object visibility to LVLMs. We present E3VQA, the first benchmark for multi-view question answering with 4K high-quality question-answer pairs grounded in synchronized ego-exo image pairs. Additionally, we propose M3CoT, a training-free prompting technique that constructs a unified scene representation by integrating scene graphs from three complementary perspectives. M3CoT enables LVLMs to reason more effectively across views, yielding consistent performance gains (4.84% for GPT-4o and 5.94% for Gemini 2.0 Flash) over a recent CoT baseline. Our extensive evaluation reveals key strengths and limitations of LVLMs in multi-view reasoning and highlights the value of leveraging both egocentric and exocentric inputs. The dataset and source code are available at https://github.com/Leeinsu1/ Towards-Comprehensive-Scene-Understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b97fd474-4956-4ed8-be5c-a083426d8a37Cited by top-tier papers1
Ask how each one uses itBuilds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Egocentric Video-Language PretrainingKevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray et al.NeurIPS 2022 · 306 citations
- DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language ModelsGe Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou et al.NeurIPS 2023 · 252 citations
Related papers
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View ScenesMohsen Gholami, Ahmad Rezaei, Zhou Weimin, Sitong Mao et al.ICLR 2026 · 67 citations
- Empowering Large Language Models with 3D Situation AwarenessZhihao Yuan, Yibo Peng, Jinke Ren, Yinghong Liao et al.CVPR 2025
- EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language ModelsSijie Cheng, Zhicheng Guo, Jingwen Wu, Kechen Fang et al.CVPR 2024
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?Liu Yang, Huiyu Duan, Ran Tao, Juntao Cheng et al.ICLR 2026 · 13 citations
- Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosQirui Chen, Shangzhe Di, Weidi XieAAAI 2025 · 35 citations
