V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
Jiancheng Pan, Runze Wang, Tianwen Qian, Mohammad Mahdi, Yanwei Fu, Xiangyang Xue, Xiaomeng Huang, Luc Van Gool, Danda Pani Paudel, Yuqian Fu
摘要
Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across different viewpoints (e.g., ego-centric and exocentric). This task poses significant challenges due to drastic viewpoint and appearance variations, making existing segmentation models, such as SAM2, non-trivial to apply directly. To address this, we present V 2 -SAM, a unified cross-view object correspondence framework that adapts SAM2 from single-view segmentation to cross-view correspondence through two complementary prompt generators. Specifically, the Cross-View Anchor Prompt Generator (V 2 -Anchor), built upon DINOv3 features, establishes geometry-aware correspondences and, for the first time, unlocks coordinate-based prompting for SAM2 in cross-view scenarios, while the Cross-View Visual Prompt Generator (V 2 -Visual) enhances appearance-guided cues via a novel visual prompt matcher that aligns ego-exo representations from both feature and structural perspectives. To effectively exploit the strengths of both prompts, we further adopt a multi-expert design and introduce a Post-hoc Cyclic Consistency Selector (PCCS) that adaptively selects the most reliable expert based on cyclic consistency. Extensive experiments validate the effectiveness of V 2 -SAM, achieving new state-of-the-art performance on Ego-Exo4D (Ego-Exo object correspondence), DAVIS-2017 (video object tracking), and HANDAL-X (robotic-ready cross-view correspondence). Codes and Models will be released at V 2 -SAM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging BenchmarkDeheng Zhang, Yuqian Fu, Runyi Yang, Yang Miao 等ICLR 2026 · 被引用 19 次
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual ReasoningKailing Li, Qi'ao Xu, Tianwen Qian, Yuqian Fu 等CVPR 2026 · 被引用 12 次
- EgoSound: Benchmarking Sound Understanding in Egocentric VideosBingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun 等CVPR 2026 · 被引用 7 次
- Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory FieldsGuangyang Wu, Youran Ding, Xinyu Che, Benyuan Sun 等CVPR 2026
- CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image PerceptionLiupeng Li, Haoqian Kang, Zhenyu Lu, Jinpeng Wang 等ICML 2026
它引用的顶会 Paper37
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li 等NeurIPS 2023 · 被引用 889 次
相关 Paper
- Robust Ego-Exo Correspondence with Long-Term MemoryYijun Hu, Bing Fan, Xin Gu, Haiqing Ren 等NeurIPS 2025 · 被引用 2 次
- O-MaMa: Learning Object Mask Matching Between Egocentric and Exocentric ViewsLorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Pérez-Yus 等ICCV 2025 · 被引用 2 次
- ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric PerspectivesYuqian Fu, Runze Wang, Bin Ren, Guolei Sun 等ICCV 2025 · 被引用 5 次
- Learning Cross-View Object Correspondence via Cycle-Consistent Mask PredictionShannan Yan, Leqi Zheng, Keyu Lv, Jingchen Ni 等CVPR 2026 · 被引用 5 次
- VGGT-Segmentor: Geometry-Enhanced Cross-View SegmentationYulu Gao, Bohao Zhang, Zongheng Tang, Jitong Liao 等CVPR 2026 · 被引用 3 次
