Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision
Xiaoshi Wu, Hadar Averbuch-Elor, Jin Sun, Noah Snavely
Abstract
The abundance and richness of Internet photos of landmarks and cities has led to significant progress in 3D vision over the past two decades, including automated 3D reconstructions of the world’s landmarks from tourist photos. However, a major source of information available for these 3D-augmented collections—namely language, e.g., from image captions—has been virtually untapped. In this work, we present WikiScenes, a new, large-scale dataset of landmark photo collections that contains descriptive text in the form of captions and hierarchical category names. WikiScenes forms a new testbed for multimodal reasoning involving images, text, and 3D geometry. We demonstrate the utility of WikiScenes for learning semantic concepts over images and 3D models. Our weakly-supervised framework connects images, 3D structure, and semantics—utilizing the strong constraints provided by 3D geometry—to associate semantic concepts to image pixels and 3D points.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Decomposing NeRF for Editing via Feature Field DistillationSosuke Kobayashi, Eiichi Matsumoto, Vincent SitzmannNeurIPS 2022 · 479 citations
- FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation ModelsJianglong Ye, Naiyan Wang, Xiaolong WangICCV 2023 · 56 citations
- Doppelgangers: Learning to Disambiguate Images of Similar StructuresRuojin Cai, Joseph Tung, Qianqian Wang, Hadar Averbuch-Elor et al.ICCV 2023 · 32 citations
- PartGlot: Learning Shape Part Segmentation from Language Reference GamesJuil Koo, Ian Huang, Panos Achlioptas, Leonidas J. Guibas et al.CVPR 2022 · 24 citations
- Combining Vision and Language Representations for Patch-based Identification of Lexico-Semantic RelationsPrince Jha, Gaël Dias, Alexis Lechervy, José G. Moreno et al.ACM MM 2022 · 5 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Shapeglot: Learning Language for Shape DifferentiationPanos Achlioptas, Leonidas J. Guibas, Noah D. Goodman, Judy Fan et al.ICCV 2019 · 86 citations
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh et al.CVPR 2020
- Single-Stage Semantic Segmentation From Image LabelsNikita Araslanov, Stefan RothCVPR 2020
- NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo CollectionsRicardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron et al.CVPR 2021
Related papers
- VisToT: Vision-Augmented Table-to-Text GenerationPrajwal Gatti, Anand Mishra, Manish Gupta, Mithun Das GuptaEMNLP 2022 · 4 citations
- Language-Assisted 3D Feature Learning for Semantic Scene UnderstandingJunbo Zhang, Guofan Fan, Guanghan Wang, Zhengyuan Su et al.AAAI 2023 · 8 citations
- Taking Language Embedded 3D Gaussian Splatting into the WildYuze Wang, Junyi Wang, Yue QiIEEE VR 2026
- Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachMathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet IscenNeurIPS 2024 · 5 citations
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited ViewsZhangquan Chen, Manyuan Zhang, Xinlei Yu, Xufang Luo et al.CVPR 2026 · 61 citations
