ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
Yuqian Fu, Runze Wang, Bin Ren, Guolei Sun, Biao Gong, Yanwei Fu, Danda Pani Paudel, Xuanjing Huang, Luc Van Gool
Abstract
Bridging the gap between ego-centric and exo-centric views has been a long-standing question in computer vision. In this paper, we focus on the emerging Ego-Exo object correspondence task, which aims to understand object relations across ego-exo perspectives through segmentation. While numerous segmentation models have been proposed, most operate on a single image (view), making them impractical for cross-view scenarios. PSALM [75], a recently proposed segmentation method, stands out as a notable exception with its demonstrated zero-shot ability on this task. However, due to the drastic viewpoint change between ego and exo, PSALM fails to accurately locate and segment objects, especially in complex backgrounds or when object appearances change significantly. To address these issues, we propose ObjectRelator, a novel approach featuring two key modules: Multimodal Condition Fusion (MCFuse) and SSL-based Cross-View Object Alignment (XObjAlign). MCFuse introduces language as an additional cue, integrating both visual masks and textual descriptions to improve object localization and prevent incorrect associations. XObjAlign enforces cross-view consistency through self-supervised alignment, enhancing robustness to object appearance variations. Extensive experiments demonstrate ObjectRelator's effectiveness on the large-scale Ego-Exo4D benchmark and HANDAL-X (an adapted dataset for cross-view segmentation) with state-of-the-art performance. Code is available at: http://yuqianfu.com/ObjectRelator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e5d3225-d180-435c-aa8f-3a518b16db8aCited by top-tier papers14
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual ReasoningKailing Li, Qi'ao Xu, Tianwen Qian, Yuqian Fu et al.CVPR 2026 · 12 citations
- Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language NavigationKailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu et al.CVPR 2026 · 10 citations
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning SegmentationZhenyu Lu, Liupeng Li, Jinpeng Wang, Yan Feng et al.ICLR 2026 · 10 citations
- Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression MethodsChenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng et al.ACL 2026 · 8 citations
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMsInsu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh et al.NeurIPS 2025 · 8 citations
Builds on29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object CorrespondenceJiancheng Pan, Runze Wang, Tianwen Qian, Mohammad Mahdi et al.CVPR 2026
- DOMR: Establishing Cross-View Segmentation via Dense Object MatchingJitong Liao, Yulu Gao, Shaofei Huang, Jialin Gao et al.ACM MM 2025
- VGGT-Segmentor: Geometry-Enhanced Cross-View SegmentationYulu Gao, Bohao Zhang, Zongheng Tang, Jitong Liao et al.CVPR 2026 · 3 citations
- O-MaMa: Learning Object Mask Matching Between Egocentric and Exocentric ViewsLorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Pérez-Yus et al.ICCV 2025 · 2 citations
- Robust Ego-Exo Correspondence with Long-Term MemoryYijun Hu, Bing Fan, Xin Gu, Haiqing Ren et al.NeurIPS 2025 · 2 citations
