Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations
Gaia Di Lorenzo, Federico Tombari, Marc Pollefeys, Daniel Barath
Abstract
Learning effective multi-modal 3D representations of objects is essential for numerous applications, such as augmented reality and robotics. Existing methods often rely on task-specific embeddings that are tailored either for semantic understanding or geometric reconstruction. As a result, these embeddings typically cannot be decoded into explicit geometry and simultaneously reused across tasks. In this paper, we propose Object-X, a versatile multi-modal object representation framework capable of encoding rich object embeddings (e.g. images, point cloud, text) and decoding them back into detailed geometric and visual reconstructions. Object-X operates by geometrically grounding the captured modalities in a 3D voxel grid and learning an unstructured embedding fusing the information from the voxels with the object attributes. The learned embedding enables 3D Gaussian Splatting-based object reconstruction, while also supporting a range of downstream tasks, including scene alignment, single-image 3D object reconstruction, and localization. Evaluations on two challenging real-world datasets demonstrate that Object-X produces high-fidelity novel-view synthesis comparable to standard 3D Gaussian Splatting, while significantly improving geometric accuracy. Moreover, Object-X achieves competitive performance with specialized methods in scene alignment and localization. Critically, our object-centric descriptors require 3-4 orders of magnitude less storage compared to traditional image- or point cloud-based approaches, establishing Object-X as a scalable and highly practical solution for multi-modal 3D scene representation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationJiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu et al.ICLR 2024 · 955 citations
- 2D Gaussian Splatting for Geometrically Accurate Radiance FieldsBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger et al.SIGGRAPH 2024 · 660 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys et al.NeurIPS 2023 · 389 citations
Related papers
- CrossOver: 3D Scene Cross-Modal AlignmentSayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath et al.CVPR 2025
- ObjectGS: Object-Aware Scene Reconstruction and Scene Understanding via Gaussian SplattingRuijie Zhu, Mulin Yu, Linning Xu, Lihan Jiang et al.ICCV 2025 · 1 citation
- UniGS: Unified Language-Image-3D Pretraining with Gaussian SplattingHaoyuan Li, Yanpeng Zhou, Tao Tang, Jifei Song et al.ICLR 2025
- ReferSplat: Referring Segmentation in 3D Gaussian SplattingShuting He, Guangquan Jie, Changshuo Wang, Yun Zhou et al.ICML 2025
- 3D Vision-Language Gaussian SplattingQucheng Peng, Benjamin Planche, Zhongpai Gao, Meng Zheng et al.ICLR 2025
