DOMR: Establishing Cross-View Segmentation via Dense Object Matching
Jitong Liao, Yulu Gao, Shaofei Huang, Jialin Gao, Jie Lei, Ronghua Liang, Si Liu
Abstract
Cross-view object correspondence involves matching objects between egocentric (first-person) and exocentric (third-person) views. It is a critical yet challenging task for visual understanding. In this work, we propose the Dense Object Matching and Refinement (DOMR) framework to establish dense object correspondences across views. The framework centers around the Dense Object Matcher (DOM) module, which jointly models multiple objects. Unlike methods that directly match individual object masks to image features, DOM leverages both positional and semantic relationships among objects to find correspondences. DOM integrates a proposal generation module with a dense matching module that jointly encodes visual, spatial, and semantic cues, explicitly constructing inter-object relationships to achieve dense matching among objects. Furthermore, we combine DOM with a mask refinement head designed to improve the completeness and accuracy of the predicted masks, forming the complete DOMR framework. Extensive evaluations on the Ego-Exo4D benchmark demonstrate that our approach achieves state-of-the-art performance with a mean IoU of 49.7% on Ego→Exo and 55.2% on Exo→Ego. These results outperform those of previous methods by 5.8% and 4.3%, respectively, validating the effectiveness of our integrated approach for cross-view understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4848631-5181-4b9f-a027-20803a3d1e91Cited by top-tier papers2
- VGGT-Segmentor: Geometry-Enhanced Cross-View SegmentationYulu Gao, Bohao Zhang, Zongheng Tang, Jitong Liao et al.CVPR 2026 · 3 citations
- V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object CorrespondenceJiancheng Pan, Runze Wang, Tianwen Qian, Mohammad Mahdi et al.CVPR 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen et al.NeurIPS 2024 · 6,113 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
Related papers
- O-MaMa: Learning Object Mask Matching Between Egocentric and Exocentric ViewsLorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Pérez-Yus et al.ICCV 2025 · 2 citations
- ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric PerspectivesYuqian Fu, Runze Wang, Bin Ren, Guolei Sun et al.ICCV 2025 · 5 citations
- Learning Cross-View Object Correspondence via Cycle-Consistent Mask PredictionShannan Yan, Leqi Zheng, Keyu Lv, Jingchen Ni et al.CVPR 2026 · 5 citations
- DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single DemoJunzhe Zhu, Yuanchen Ju, Junyi Zhang, Muhan Wang et al.ICLR 2025
- Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video RepresentationsJungin Park, Jiyoung Lee, Kwanghoon SohnCVPR 2025
