Improving Scene Graph Classification by Exploiting Knowledge from Texts
Sahand Sharifzadeh, Sina Moayed Baharlou, Martin Schmitt, Hinrich Schütze, Volker Tresp
摘要
Training scene graph classification models requires a large amount of annotated image data. Meanwhile, scene graphs represent relational knowledge that can be modeled with symbolic data from texts or knowledge graphs. While image annotation demands extensive labor, collecting textual descriptions of natural scenes requires less effort. In this work, we investigate whether textual scene descriptions can substitute for annotated image data. To this end, we employ a scene graph classification framework that is trained not only from annotated images but also from symbolic data. In our architecture, the symbolic entities are first mapped to their correspondent image-grounded representations and then fed into the relational reasoning pipeline. Even though a structured form of knowledge, such as the form in knowledge graphs, is not always available, we can generate it from unstructured texts using a transformer-based language model. We show that by fine-tuning the classification pipeline with the extracted knowledge from texts, we can achieve 8x more accurate results in scene graph classification, 3x in object classification, and 1.5x in predicate classification, compared to the supervised baselines with only 1% of the annotated images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Multi-Prototype Space Learning for Commonsense-Based Scene Graph GenerationLianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu 等AAAI 2024 · 被引用 11 次
- Weakly-Supervised Learning of Visual Relations in Multimodal PretrainingEmanuele Bugliarello, Aida Nematzadeh, Lisa Anne HendricksEMNLP 2023 · 被引用 1 次
- DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph RefinementShaoqing Lin, Chong Teng, Fei Li, Donghong Ji 等EMNLP 2025
- Universal Scene Graph GenerationShengqiong Wu, Hao Fei, Tat-Seng ChuaCVPR 2025
它引用的顶会 Paper4
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He 等ICCV 2019 · 被引用 165 次
- Classification by Attention: Scene Graph Classification with Prior KnowledgeSahand Sharifzadeh, Sina Moayed Baharlou, Volker TrespAAAI 2021 · 被引用 61 次
- An Unsupervised Joint System for Text Generation from Knowledge Graphs and Semantic ParsingMartin Schmitt, Sahand Sharifzadeh, Volker Tresp, Hinrich SchützeEMNLP 2020 · 被引用 2 次
相关 Paper
- TextPSG: Panoptic Scene Graph Generation from Textual DescriptionsChengyang Zhao, Yikang Shen, Zhenfang Chen, Mingyu Ding 等ICCV 2023 · 被引用 24 次
- Scalable Theory-Driven Regularization of Scene Graph Generation ModelsDavide Buffelli, Efthymia TsamouraAAAI 2023 · 被引用 4 次
- Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene GraphsRoei Herzig, Alon Mendelson, Leonid Karlinsky, Assaf Arbelle 等EMNLP 2023 · 被引用 15 次
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu 等ICCV 2021 · 被引用 88 次
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen 等CVPR 2024
