Visual Intention Grounding for Egocentric Assistants
Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li, Arjun R. Akula, Angela Yao
Abstract
Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts - inputs are egocentric, and objects may be referred to implicitly through needs and intentions. To bridge this gap, we introduce EgoIntention, the first dataset for egocentric visual intention grounding. EgoIntention challenges multimodal LLMs to 1) understand and ignore unintended contextual objects and 2) reason about uncommon object functionalities. Benchmark results show that current models misidentify context objects and lack affordance understanding in egocentric views. We also propose Reason-to-Ground (RoG) instruction tuning; it enables hybrid training with normal descriptions and egocentric intentions with a chained intention reasoning and object grounding mechanism. RoG significantly outperforms naive finetuning and hybrid training on EgoIntention, while maintaining or slightly improving naive description grounding. This advancement enables unified visual grounding for egocentric and exocentric visual inputs while handling explicit object queries and implicit human intentions. Our code and model are available at https://github.com/pengzhansun/EgoIntention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Ego-Grounding for Personalized Question-Answering in Egocentric VideosJunbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela YaoCVPR 2026 · 7 citations
- Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object CategoriesYicong Li, Yiyang Chen, Zhenyuan Ma, Junbin Xiao et al.ICCV 2025 · 3 citations
- EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive HierarchyJinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan et al.CVPR 2026 · 2 citations
- Learning Scene Coordinate Reconstruction from Unposed Images via Pose Graph OptimizationTze Ho Elden Tse, Jizong Peng, Angela YaoCVPR 2026
Builds on17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Ferret: Refer and Ground Anything Anywhere at Any GranularityHaoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du et al.ICLR 2024 · 515 citations
Related papers
- Intent3D: 3D Object Detection in RGB-D Scans Based on Human IntentionWeitai Kang, Mengxue Qu, Jyoti Kini, Yunchao Wei et al.ICLR 2025
- TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free TuningAritra Bhowmik, Mohammad Mahdi Derakhshani, Dennis C. Koelma, Yuki M. Asano et al.ICCV 2025 · 2 citations
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai et al.AAAI 2026
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoTBaoqi Pei, Yifei Huang, Jilan Xu, Yuping He et al.NeurIPS 2025 · 21 citations
- Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction TuningJi Hyeok Jung, Eun Tae Kim, Seo Yeon Kim, Joo Ho Lee et al.CVPR 2025
