Lune

ICLR2025Top-tier venue

Duoduo CLIP: Efficient 3D Understanding with Multi-View Images

Han-Hung Lee, Yiming Zhang, Angel X. Chang

2025Year
6Top-tier citations

Abstract

We introduce Duoduo CLIP, a model for 3D representation learning that learns shape encodings from multi-view images instead of point clouds. The choice of multi-view images allows us to leverage 2D priors from off-the-shelf CLIP models to facilitate fine-tuning with 3D data. Our approach not only shows better generalization compared to existing point cloud methods, but also reduces GPU requirements and training time. In addition, the model is modified with cross-view attention to leverage information across multiple frames of the object which further boosts performance. Notably, our model is permutation invariant to the order of multi-view images while being pose-free. Compared to the current SOTA point cloud method that requires 480 A100 hours to train 1 billion model parameters we only require 57 A5000 hours and 87 million parameters. Multi-view images also provide more flexibility including being able to encode objects with a variable number of images, and performance scales when more views are used. In contrast, point cloud based methods require an entire scan or model of the object. We showcase this flexibility with benchmarks from images of real-world objects. Our model also achieves better performance in more fine-grained text to shape retrieval, demonstrating better text-and-shape alignment than point cloud based models. RELATED WORK There is much work on 3D representation learning using supervised training (Qi et al., 2017; Qian et al., 2022) or using self-supervised learning (with reconstruction loss (Chen & Zhang, 2019), METHOD Our Duoduo CLIP learns a shape representation by using contrastive learning to encode multi-view images into a pre-aligned text-image space. Our contrastive learning framework is similar to previous works (

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 91c4d5a8-20c2-4b81-ad45-e8012c253665

Cited by top-tier papers6

Ask how each one uses it

Builds on37

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines