Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction
Shannan Yan, Leqi Zheng, Keyu Lv, Jingchen Ni, Hongyang Wei, Jiajun Zhang, Guangting Wang, Jing LYU, Chun Yuan, Fengyun Rao
摘要
We study the task of establishing object-level visual correspondence across different viewpoints in videos, focusing on the challenging egocentric-to-exocentric and exocentricto-egocentric scenarios. We propose a simple yet effective framework based on conditional binary segmentation, where an object query mask is encoded into a latent representation to guide the localization of the corresponding object in a target video. To encourage robust, view-invariant representations, we introduce a cycle-consistency training objective: the predicted mask in the target view is projected back to the source view to reconstruct the original query mask. This bidirectional constraint provides a strong self-supervisory signal without requiring ground-truth annotations and enables test-time training (TTT) at inference. Experiments on the Ego-Exo4D and HANDAL-X benchmarks demonstrate the effectiveness of our optimization objective and TTT strategy, achieving state-of-the-art performance. The code is available at https://github. com/shannany0606/CCMP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context LearningTianci Luo, Haohao Pan, Jinpeng Wang, Niu Lian 等CVPR 2026
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo 等ACL 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object CorrespondenceJiancheng Pan, Runze Wang, Tianwen Qian, Mohammad Mahdi 等CVPR 2026
- Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningZixu Zhao, Yueming Jin, Pheng-Ann HengICCV 2021 · 被引用 23 次
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 被引用 356 次
- Self-Supervised Cross-View Correspondence with Predictive Cycle ConsistencyAlan Baade, Changan ChenCVPR 2025
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 被引用 64 次
