Unsupervised Vision-Language Grammar Induction with Shared Structure Modeling
Bo Wan, Wenjuan Han, Zilong Zheng, Tinne Tuytelaars
Abstract
We introduce a new task, unsupervised vision-language (VL) grammar induction. Given an image-caption pair, the goal is to extract a shared hierarchical structure for both image and language simultaneously. We argue that such structured output, grounded in both modalities, is a clear step towards the high-level understanding of multimodal information. Besides challenges existing in conventional visually grounded grammar induction tasks, VL grammar induction requires a model to capture contextual semantics and perform a fine-grained alignment. To address these challenges, we propose a novel method, CLIORA, which constructs a shared vision-language constituency tree structure with context-dependent semantics for all possible phrases in different levels of the tree. It computes a matching score between each constituent and image region, trained via contrastive learning. It integrates two levels of fusion, namely at feature-level and at score-level, so as to allow fine-grained alignment. We introduce a new evaluation metric for VL grammar induction, CCRA, and show a 3.3% improvement over a strong baseline on Flickr30k Entities. We also evaluate our model via two derived tasks, i.e., language grammar induction and phrase grounding, and improve over the state-of-the-art for both.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9991adde-5cc9-4f41-9130-fb2c7c7a5205Cited by top-tier papers5
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
- Activity Grammars for Temporal Action SegmentationDayoung Gong, Joonseok Lee, Deunsol Jung, Suha Kwak et al.NeurIPS 2023 · 17 citations
- HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware AttentionShijie Geng, Jianbo Yuan, Yu Tian, Yuxiao Chen et al.ICLR 2023 · 11 citations
- Cross-modal Attention Congruence Regularization for Vision-Language Relation AlignmentRohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov et al.ACL 2023 · 3 citations
- Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at ScaleXiang Hu, Pengyu Ji, Qingyang Zhu, Wei Wu et al.ACL 2024 · 1 citation
Related papers
- VLGrammar: Grounded Grammar Induction of Vision and LanguageYining Hong, Qing Li, Song-Chun Zhu, Siyuan HuangICCV 2021 · 28 citations
- Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency RelationshipsChao Lou, Wenjuan Han, Yuhuan Lin, Zilong ZhengCVPR 2022 · 9 citations
- Contrastive Learning with Expectation-Maximization for Weakly Supervised Phrase GroundingKeqin Chen, Richong Zhang, Samuel Mensah, Yongyi MaoEMNLP 2022 · 3 citations
- Visually Grounded Compound PCFGsYanpeng Zhao, Ivan TitovEMNLP 2020 · 35 citations
- HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual GroundingLinhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang et al.ACM MM 2024 · 28 citations
