Shapeglot: Learning Language for Shape Differentiation
Panos Achlioptas, Leonidas J. Guibas, Noah D. Goodman, Judy Fan, Robert X. D. Hawkins
摘要
In this work we explore how fine-grained differences between the shapes of common objects are expressed in language, grounded on 2D and/or 3D object representations. We first build a large scale, carefully controlled dataset of human utterances each of which refers to a 2D rendering of a 3D CAD model so as to distinguish it from a set of shape-wise similar alternatives. Using this dataset, we develop neural language understanding (listening) and production (speaking) models that vary in their grounding (pure 3D forms via point-clouds vs. rendered 2D images), the degree of pragmatic reasoning captured (e.g. speakers that reason about a listener or not), and the neural architecture (e.g. with or without attention). We find models that perform well with both synthetic and human partners, and with held out utterances and objects. We also find that these models are capable of zero-shot transfer learning to novel object classes (e.g. transfer from training on chairs to testing on lamps), as well as to real-world images drawn from furniture catalogs. Lesion studies indicate that the neural listeners depend heavily on part-related words and associate these words correctly with visual parts of objects (without any explicit supervision on such parts), and that transfer to novel classes is most successful when known part-related words are available. This work illustrates a practical approach to language grounding, and provides a novel case study in the relationship between object shape and linguistic structure when it comes to object differentiation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 被引用 191 次
- 3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion ModelsBiao Zhang, Jiapeng Tang, Matthias Nießner, Peter WonkaSIGGRAPH 2023 · 被引用 172 次
- Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsYiming Zhang, ZeMing Gong, Angel X. ChangICCV 2023 · 被引用 157 次
- ShapeCrafter: A Recursive Text-Conditioned 3D Shape Generation ModelRao Fu, Xiao Zhan, Yiwen Chen, Daniel Ritchie 等NeurIPS 2022 · 被引用 98 次
- SALAD: Part-Level Latent Diffusion for 3D Shape Generation and ManipulationJuil Koo, Seungwoo Yoo, Minh Hieu Nguyen, Minhyuk SungICCV 2023 · 被引用 79 次
它引用的顶会 Paper1
相关 Paper
- PartGlot: Learning Shape Part Segmentation from Language Reference GamesJuil Koo, Ian Huang, Panos Achlioptas, Leonidas J. Guibas 等CVPR 2022 · 被引用 24 次
- ShapeTalk: A Language Dataset and Framework for 3D Shape Edits and DeformationsPanos Achlioptas, Ian Huang, Minhyuk Sung, Sergey Tulyakov 等CVPR 2023
- Language Conditioned Spatial Relation Reasoning for 3D Object GroundingShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid 等NeurIPS 2022 · 被引用 173 次
- Advancing 3D Object Grounding Beyond a Single 3D SceneWencan Huang, Daizong Liu, Wei HuACM MM 2024 · 被引用 9 次
- Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set RelationshipsSebastian Koch, Narunas Vaskevicius, Mirco Colosi, Pedro Hermosilla 等CVPR 2024
