Linguistic Structures As Weak Supervision for Visual Scene Graph Generation
Keren Ye, Adriana Kovashka
Abstract
Prior work in scene graph generation requires categorical supervision at the level of triplets-subjects and objects, and predicates that relate them, either with or without bounding box information. However, scene graph generation is a holistic task: thus holistic, contextual supervision should intuitively improve performance. In this work, we explore how linguistic structures in captions can benefit scene graph generation. Our method captures the information provided in captions about relations between individual triplets, and context for subjects and objects (e.g. visual properties are mentioned). Captions are a weaker type of supervision than triplets since the alignment between the exhaustive list of human-annotated subjects and objects in triplets, and the nouns in captions, is weak. However, given the large and diverse sources of multimodal data on the web (e.g. blog posts with images and captions), linguistic supervision is more scalable than crowdsourced triplets. We show extensive experimental comparisons against prior methods which leverage instance-and image-level supervision, and ablate our method to show the impact of leveraging phrasal and sequential context, and techniques to improve localization of subjects and objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu et al.ICCV 2021 · 88 citations
- A Simple Baseline for Weakly-Supervised Scene Graph GenerationJing Shi, Yiwu Zhong, Ning Xu, Yin Li et al.ICCV 2021 · 34 citations
- Scene Graph Generation with Role-Playing Large Language ModelsGuikun Chen, Jin Li, Wenguan WangNeurIPS 2024 · 33 citations
- TextPSG: Panoptic Scene Graph Generation from Textual DescriptionsChengyang Zhao, Yikang Shen, Zhenfang Chen, Mingyu Ding et al.ICCV 2023 · 24 citations
- Integrating Object-aware and Interaction-aware Knowledge for Weakly Supervised Scene Graph GenerationXingchen Li, Long Chen, Wenbo Ma, Yi Yang et al.ACM MM 2022 · 22 citations
Builds on8
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Cap2Det: Learning to Amplify Weak Caption Supervision for Object DetectionKeren Ye, Mingda Zhang, Adriana Kovashka, Wei Li et al.ICCV 2019 · 61 citations
- VirTex: Learning Visual Representations From Textual AnnotationsKaran Desai, Justin JohnsonCVPR 2021
- GPS-Net: Graph Property Sensing Network for Scene Graph GenerationXin Lin, Changxing Ding, Jinquan Zeng, Dacheng TaoCVPR 2020
- Hypergraph Attention Networks for Multimodal LearningEun-Sol Kim, Woo-Young Kang, Kyoung-Woon On, Yu-Jung Heo et al.CVPR 2020
Related papers
- LLM4SGG: Large Language Models for Weakly Supervised Scene Graph GenerationKibum Kim, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In et al.CVPR 2024
- Weakly-Supervised Learning of Visual Relations in Multimodal PretrainingEmanuele Bugliarello, Aida Nematzadeh, Lisa Anne HendricksEMNLP 2023 · 1 citation
- In Defense of Scene Graphs for Image CaptioningKien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha et al.ICCV 2021 · 55 citations
- Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image CaptioningWeizhi Nie, Jiesi Li, Ning Xu, An-An Liu et al.ACM MM 2021 · 9 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
