One-shot Scene Graph Generation
Yuyu Guo, Jingkuan Song, Lianli Gao, Heng Tao Shen
摘要
As a structured representation of the image content, the visual scene graph (visual relationship) acts as a bridge between computer vision and natural language processing. Existing models on the scene graph generation task notoriously require tens or hundreds of labeled samples. By contrast, human beings can learn visual relationships from a few or even one example. Inspired by this, we design a task named One-Shot Scene Graph Generation, where each relationship triplet (e.g., "dog-has-head'') comes from only one labeled example. The key insight is that rather than learning from scratch, one can utilize rich prior knowledge. In this paper, we propose Multiple Structured Knowledge (Relational Knowledge and Commonsense Knowledge) for the one-shot scene graph generation task. Specifically, the Relational Knowledge represents the prior knowledge of relationships between entities extracted from the visual content, e.g., the visual relationships "standing in'', "sitting in'', and "lying in'' may exist between "dog'' and "yard'', while the Commonsense Knowledge encodes "sense-making'' knowledge like "dog can guard yard''. By organizing these two kinds of knowledge in a graph structure, Graph Convolution Networks (GCNs) are used to extract knowledge-embedded semantic features of the entities. Besides, instead of extracting isolated visual features from each entity generated by Faster R-CNN, we utilize an Instance Relation Transformer encoder to fully explore their context information. Based on a constructed one-shot dataset, the experimental results show that our method significantly outperforms existing state-of-the-art methods by a large margin. Ablation studies also verify the effectiveness of the Instance Relation Transformer encoder and the Multiple Structured Knowledge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu 等ICCV 2021 · 被引用 96 次
- Fine-Grained Predicates Learning for Scene Graph GenerationXinyu Lyu, Lianli Gao, Yuyu Guo, Zhou Zhao 等CVPR 2022 · 被引用 48 次
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 被引用 47 次
- HGOE: Hybrid External and Internal Graph Outlier Exposure for Graph Out-of-Distribution DetectionJunwei He, Qianqian Xu, Yangbangyan Jiang, Zitai Wang 等ACM MM 2024 · 被引用 4 次
- Prototype-Based Embedding Network for Scene Graph GenerationChaofan Zheng, Xinyu Lyu, Lianli Gao, Bo Dai 等CVPR 2023
它引用的顶会 Paper3
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He 等ICCV 2019 · 被引用 165 次
- Music Gesture for Visual Sound SeparationChuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum 等CVPR 2020
- Universal Weighting Metric Learning for Cross-Modal MatchingJiwei Wei, Xing Xu, Yang Yang, Yanli Ji 等CVPR 2020
相关 Paper
- One-Shot Learning for Long-Tail Visual Relation DetectionWeitao Wang, Meng Wang, Sen Wang, Guodong Long 等AAAI 2020 · 被引用 20 次
- SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense ReasoningZhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian 等AAAI 2022 · 被引用 40 次
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 被引用 2 次
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang 等AAAI 2020 · 被引用 75 次
- Visual Distant Supervision for Scene Graph GenerationYuan Yao, Ao Zhang, Xu Han, Mengdi Li 等ICCV 2021 · 被引用 41 次
