Learning Geometric-Aware Properties in 2D Representation Using Lightweight CAD Models, or Zero Real 3D Pairs
Pattaramanee Arsomngern, Sarana Nutanong, Supasorn Suwajanakorn
摘要
Cross-modal training using 2D-3D paired datasets, such as those containing multi-view images and 3D scene scans, presents an effective way to enhance 2D scene understanding by introducing geometric and view-invariance priors into 2D features. However, the need for large-scale scene datasets can impede scalability and further improvements. This paper explores an alternative learning method by leveraging a lightweight and publicly available type of 3D data in the form of CAD models. We construct a 3D space with geometric-aware alignment where the similarity in this space reflects the geometric similarity of CAD models based on the Chamfer distance. The acquired geometricaware properties are then induced into 2D features, which boost performance on downstream tasks more effectively than existing RGB-CAD approaches. Our technique is not limited to paired RGB-CAD datasets. By training exclusively on pseudo pairs generated from CAD-based reconstruction methods, we enhance the performance of SOTA 2D pre-trained models that use ResNet-50 or ViT-B backbones on various 2D understanding tasks. We also achieve comparable results to SOTA methods trained on scene scans on four tasks in NYUv2, SUNRGB-D, indoor ADE20k, and indoor/outdoor COCO, despite using lightweight CAD models or pseudo data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Perspective from a Broader Context: Can Room Style Knowledge Help Visual Floorplan Localization?Bolei Chen, Shengsheng Yan, Yongzheng Cui, Jiaxu Kang 等AAAI 2026 · 被引用 1 次
- Perspective from a Higher Dimension: Can 3D Geometric Priors Help Visual Floorplan Localization?Bolei Chen, Jiaxu Kang, Haonan Yang, Ping Zhong 等ACM MM 2025
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Patch2CAD: Patchwise Embedding Learning for In-the-Wild Shape Retrieval from a Single ImageWeicheng Kuo, Anelia Angelova, Tsung-Yi Lin, Angela DaiICCV 2021 · 被引用 42 次
- Pri3D: Can 3D Priors Help 2D Representation Learning?Ji Hou, Saining Xie, Benjamin Graham, Angela Dai 等ICCV 2021 · 被引用 94 次
- CrossOver: 3D Scene Cross-Modal AlignmentSayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath 等CVPR 2025
- Mask3D: Pretraining 2D Vision Transformers by Learning Masked 3D PriorsJi Hou, Xiaoliang Dai, Zijian He, Angela Dai 等CVPR 2023
- RandomRooms: Unsupervised Pre-training from Synthetic Shapes and Randomized Layouts for 3D Object DetectionYongming Rao, Benlin Liu, Yi Wei, Jiwen Lu 等ICCV 2021 · 被引用 58 次
