Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
Yang You, Yixin Li, Congyue Deng, Yue Wang, Leonidas J. Guibas
摘要
Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D awareness of ViT-based models. We begin by systematically assessing their ability to learn 3D equivariant features, specifically examining the consistency of semantic embeddings across different viewpoints. Our findings indicate that improved 3D equivariance leads to better performance on various downstream tasks, including pose estimation, tracking, and semantic transfer. Building on this insight, we propose a simple yet effective finetuning strategy based on 3D correspondences, which significantly enhances the 3D correspondence understanding of existing vision models. Remarkably, finetuning on a single object for one iteration results in substantial gains. Our code is available at https://github.com/qq456cvb/3DCorrEnhance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Spatial Mental Modeling from Limited ViewsQineng Wang, Baiqiao Yin, Pingyue Zhang, Jianshu Zhang 等ICLR 2026 · 被引用 92 次
- 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene UnderstandingXiaohu Huang, Jingjing Wu, Qunyi Xie, Kai HanNeurIPS 2025 · 被引用 11 次
- Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware DistillationDavid Shavin, Sagie BenaimICLR 2026 · 被引用 2 次
- Cortical Policy: A Dual-Stream View Transformer for Robotic ManipulationXuening Zhang, Qi Lv, Xiang Deng, Miao Zhang 等ICLR 2026 · 被引用 1 次
- WarpHE4D: Dense 4D Head Map Toward Full Head ReconstructionJongseob Yun, Yong-Hoon Kwon, Min-Gyu Park, Ju-Mi Kang 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa 等ICCV 2023 · 被引用 620 次
相关 Paper
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- ExtPose: Robust and Coherent Pose Estimation by Extending ViTsRongyu Chen, Li'an Zhuo, Linlin Yang, Qi Wang 等ICML 2025
- SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth EstimationZiyan He, Qiudan Zhang, Lin Ma, Xu WangCVPR 2026
- Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceJinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 被引用 18 次
- Uni3D: Exploring Unified 3D Representation at ScaleJunsheng Zhou, Jinsheng Wang, Baorui Ma, Yu-Shen Liu 等ICLR 2024 · 被引用 207 次
