One-shot Scene Graph Generation
Yuyu Guo, Jingkuan Song, Lianli Gao, Heng Tao Shen
Abstract
As a structured representation of the image content, the visual scene graph (visual relationship) acts as a bridge between computer vision and natural language processing. Existing models on the scene graph generation task notoriously require tens or hundreds of labeled samples. By contrast, human beings can learn visual relationships from a few or even one example. Inspired by this, we design a task named One-Shot Scene Graph Generation, where each relationship triplet (e.g., "dog-has-head'') comes from only one labeled example. The key insight is that rather than learning from scratch, one can utilize rich prior knowledge. In this paper, we propose Multiple Structured Knowledge (Relational Knowledge and Commonsense Knowledge) for the one-shot scene graph generation task. Specifically, the Relational Knowledge represents the prior knowledge of relationships between entities extracted from the visual content, e.g., the visual relationships "standing in'', "sitting in'', and "lying in'' may exist between "dog'' and "yard'', while the Commonsense Knowledge encodes "sense-making'' knowledge like "dog can guard yard''. By organizing these two kinds of knowledge in a graph structure, Graph Convolution Networks (GCNs) are used to extract knowledge-embedded semantic features of the entities. Besides, instead of extracting isolated visual features from each entity generated by Faster R-CNN, we utilize an Instance Relation Transformer encoder to fully explore their context information. Based on a constructed one-shot dataset, the experimental results show that our method significantly outperforms existing state-of-the-art methods by a large margin. Ablation studies also verify the effectiveness of the Instance Relation Transformer encoder and the Multiple Structured Knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu et al.ICCV 2021 · 96 citations
- Fine-Grained Predicates Learning for Scene Graph GenerationXinyu Lyu, Lianli Gao, Yuyu Guo, Zhou Zhao et al.CVPR 2022 · 48 citations
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 47 citations
- HGOE: Hybrid External and Internal Graph Outlier Exposure for Graph Out-of-Distribution DetectionJunwei He, Qianqian Xu, Yangbangyan Jiang, Zitai Wang et al.ACM MM 2024 · 4 citations
- Prototype-Based Embedding Network for Scene Graph GenerationChaofan Zheng, Xinyu Lyu, Lianli Gao, Bo Dai et al.CVPR 2023
Builds on3
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He et al.ICCV 2019 · 165 citations
- Music Gesture for Visual Sound SeparationChuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum et al.CVPR 2020
- Universal Weighting Metric Learning for Cross-Modal MatchingJiwei Wei, Xing Xu, Yang Yang, Yanli Ji et al.CVPR 2020
Related papers
- One-Shot Learning for Long-Tail Visual Relation DetectionWeitao Wang, Meng Wang, Sen Wang, Guodong Long et al.AAAI 2020 · 20 citations
- SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense ReasoningZhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian et al.AAAI 2022 · 40 citations
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 2 citations
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang et al.AAAI 2020 · 75 citations
- Visual Distant Supervision for Scene Graph GenerationYuan Yao, Ao Zhang, Xu Han, Mengdi Li et al.ICCV 2021 · 41 citations
