VLGrammar: Grounded Grammar Induction of Vision and Language
Yining Hong, Qing Li, Song-Chun Zhu, Siyuan Huang
摘要
Cognitive grammar suggests that the acquisition of language grammar is grounded within visual structures. While grammar is an essential representation of natural language, it also exists ubiquitously in vision to represent the hierarchical part-whole structure. In this work, we study grounded grammar induction of vision and language in a joint learning framework. Specifically, we present VLGrammar, a method that uses compound probabilistic context-free grammars (compound PCFGs) to induce the language grammar and the image grammar simultaneously. We propose a novel contrastive learning framework to guide the joint learning of both modules. To provide a benchmark for the grounded grammar induction task, we collect a large-scale dataset, PARTIT, which contains human-written sentences that describe part-level semantics for 3D objects. Experiments on the PARTIT dataset show that VLGrammar outperforms all baselines in image grammar induction and language grammar induction. The learned VLGrammar naturally benefits related downstream tasks. Specifically, it improves the image unsupervised clustering accuracy by 30%, and performs well in image retrieval and text retrieval. Notably, the induced grammar shows superior generalizability by easily generalizing to unseen categories. Code and pre-trained models are released at https://github.com/evelinehong/VLGrammar.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- 3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentZiyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng 等ICCV 2023 · 被引用 247 次
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu 等NeurIPS 2022 · 被引用 207 次
- PTR: A Benchmark for Part-based Conceptual, Relational, and Physical ReasoningYining Hong, Li Yi, Josh Tenenbaum, Antonio Torralba 等NeurIPS 2021 · 被引用 46 次
- Sequence-to-Sequence Learning with Latent Neural GrammarsYoon KimNeurIPS 2021 · 被引用 44 次
- PartGlot: Learning Shape Part Segmentation from Language Reference GamesJuil Koo, Ian Huang, Panos Achlioptas, Leonidas J. Guibas 等CVPR 2022 · 被引用 24 次
它引用的顶会 Paper4
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Closed Loop Neural-Symbolic Learning via Integrating Neural Perception, Grammar Parsing, and Symbolic ReasoningQing Li, Siyuan Huang, Yining Hong, Yixin Chen 等ICML 2020 · 被引用 93 次
- Visually Grounded Compound PCFGsYanpeng Zhao, Ivan TitovEMNLP 2020 · 被引用 35 次
相关 Paper
- Unsupervised Vision-Language Grammar Induction with Shared Structure ModelingBo Wan, Wenjuan Han, Zilong Zheng, Tinne TuytelaarsICLR 2022 · 被引用 19 次
- Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency RelationshipsChao Lou, Wenjuan Han, Yuhuan Lin, Zilong ZhengCVPR 2022 · 被引用 9 次
- Universal 3D Shape Matching via Coarse-to-Fine Language GuidanceQinfeng Xiao, Guofeng Mei, Bo Yang, Zhang Liying 等CVPR 2026 · 被引用 1 次
- Activity Grammars for Temporal Action SegmentationDayoung Gong, Joonseok Lee, Deunsol Jung, Suha Kwak 等NeurIPS 2023 · 被引用 17 次
- Object-centric binding in Contrastive Language-Image PretrainingRim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal 等NeurIPS 2025 · 被引用 14 次
