Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency Relationships
Chao Lou, Wenjuan Han, Yuhuan Lin, Zilong Zheng
摘要
Understanding realistic visual scene images together with language descriptions is a fundamental task towards generic visual understanding. Previous works have shown compelling comprehensive results by building hierarchical structures for visual scenes (e.g., scene graphs) and natural languages (e.g., dependency trees), individually. However, how to construct a joint vision-language (VL) structure has barely been investigated. More challenging but worthwhile, we introduce a new task that targets on inducing such a joint VL structure in an unsupervised manner. Our goal is to bridge the visual scene graphs and linguistic dependency trees seamlessly. Due to the lack of VL structural data, we start by building a new dataset VLParse. Rather than using labor-intensive labeling from scratch, we propose an automatic alignment procedure to produce coarse structures followed by human refinement to produce high-quality ones. Moreover, we benchmark our dataset by proposing a contrastive learning (CL)-based framework VLGAE, short for Vision-Language Graph Autoencoder. Our model obtains superior performance on two derived tasks, i.e., language grammar induction and VL phrase grounding. Ablations show the effectiveness of both visual cues and dependency relationships on fine-grained VL structure construction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 被引用 97 次
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani 等ICLR 2023 · 被引用 70 次
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada 等ICCV 2023 · 被引用 42 次
- CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual GroundingEslam Mohamed Bakr, Mohamed Ayman, Mahmoud Ahmed, Habib Slim 等ICLR 2024 · 被引用 16 次
- Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual SceneShengqiong Wu, Hao Fei, Jingkang Yang, Xiangtai Li 等CVPR 2025
它引用的顶会 Paper8
- Efficient Second-Order TreeCRF for Neural Dependency ParsingYu Zhang, Zhenghua Li, Min ZhangACL 2020 · 被引用 90 次
- Visually Grounded Compound PCFGsYanpeng Zhao, Ivan TitovEMNLP 2020 · 被引用 35 次
- A Simple Baseline for Weakly-Supervised Scene Graph GenerationJing Shi, Yiwu Zhong, Ning Xu, Yin Li 等ICCV 2021 · 被引用 34 次
- VLGrammar: Grounded Grammar Induction of Vision and LanguageYining Hong, Qing Li, Song-Chun Zhu, Siyuan HuangICCV 2021 · 被引用 28 次
- Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive AutoencodersAndrew Drozdov, Subendhu Rongali, Yi-Pei Chen, Tim O'Gorman 等EMNLP 2020 · 被引用 27 次
相关 Paper
- Unsupervised Vision-Language Grammar Induction with Shared Structure ModelingBo Wan, Wenjuan Han, Zilong Zheng, Tinne TuytelaarsICLR 2022 · 被引用 19 次
- Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language CompositionalityHarman Singh, Pengchuan Zhang, Qifan Wang, Mengjiao Wang 等EMNLP 2023 · 被引用 10 次
- PARSE: Part-Aware Relational Spatial ModelingYinuo Bai, Peijun Xu, Kuixiang Shao, Yuyang Jiao 等CVPR 2026
- Multimodal Contextualized Semantic Parsing from SpeechJordan Voas, David Harwath, Raymond MooneyACL 2024
- Extending Phrase Grounding with Pronouns in Visual DialoguesPanzhong Lu, Xin Zhang, Meishan Zhang, Min ZhangEMNLP 2022 · 被引用 5 次
