Multi-Prototype Space Learning for Commonsense-Based Scene Graph Generation
Lianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu, Yang Li, Yuan Xie, Changbo Wang, Gaoqi He
Abstract
In the domain of scene graph generation, modeling commonsense as a single-prototype representation has been typically employed to facilitate the recognition of infrequent predicates. However, a fundamental challenge lies in the large intra-class variations of the visual appearance of predicates, resulting in subclasses within a predicate class. Such a challenge typically leads to the problem of misclassifying diverse predicates due to the rough predicate space clustering. In this paper, inspired by cognitive science, we maintain multi-prototype representations for each predicate class, which can accurately find the multiple class centers of the predicate space. Technically, we propose a novel multi-prototype learning framework consisting of three main steps: prototype-predicate matching, prototype updating, and prototype space optimization. We first design a triple-level optimal transport to match each predicate feature within the same class to a specific prototype. In addition, the prototypes are updated using momentum updating to find the class centers according to the matching results. Finally, we enhance the inter-class separability of the prototype space through iterations of the inter-class separability loss and intra-class compactness loss. Extensive evaluations demonstrate that our approach significantly outperforms state-of-the-art methods on the Visual Genome dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 458e40fd-2eeb-4a68-a5e4-51915d81a77eCited by top-tier papers3
- Prototype-Guided Multimodal Relation Extraction based on Entity AttributesZefan Zhang, Weiqi Zhang, Yanhui Li, Tian BaiAAAI 2025 · 8 citations
- Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationChanggu Chen, Junwei Shu, Gaoqi He, Changbo Wang et al.AAAI 2025 · 1 citation
- Learning Context-Conditioned Predicate Semantics via Prototype FeedbackNamGyu Jung, Chang ChoiICML 2026
Builds on20
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph GenerationLin Li, Long Chen, Yifeng Huang, Zhimeng Zhang et al.CVPR 2022 · 103 citations
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu et al.ICCV 2021 · 88 citations
- ReFormer: The Relational Transformer for Image CaptioningXuewen Yang, Yingru Liu, Xin WangACM MM 2022 · 70 citations
- Cross-Domain Imitation Learning via Optimal TransportArnaud Fickinger, Samuel Cohen, Stuart Russell, Brandon AmosICLR 2022 · 65 citations
Related papers
- Prototype-Based Embedding Network for Scene Graph GenerationChaofan Zheng, Xinyu Lyu, Lianli Gao, Bo Dai et al.CVPR 2023
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu et al.ICCV 2021 · 96 citations
- Classification by Attention: Scene Graph Classification with Prior KnowledgeSahand Sharifzadeh, Sina Moayed Baharlou, Volker TrespAAAI 2021 · 61 citations
- Synergetic Prototype Learning Network for Unbiased Scene Graph GenerationRuonan Zhang, Ziwei Shang, Fengjuan Wang, Zhaoqilin Yang et al.ACM MM 2024 · 5 citations
- Beware of Overcorrection: Scene-induced Commonsense Graph for Scene Graph GenerationLianggangxu Chen, Jiale Lu, Youqi Song, Changbo Wang et al.ACM MM 2023 · 5 citations
