Leveraging Predicate and Triplet Learning for Scene Graph Generation
Jiankai Li, Yunhong Wang, Xiefan Guo, Ruijie Yang, Weixin Li
Abstract
Scene Graph Generation (SGG) aims to identify entities and predict the relationship triplets <subject, predicate, object> in visual scenes. Given the prevalence of large visual variations of subject-object pairs even in the same predicate, it can be quite challenging to model and refine predicate representations directly across such pairs, which is however a common strategy adopted by most existing SGG methods. We observe that visual variations within the identical triplet are relatively small and certain relation cues are shared in the same type of triplet, which can potentially facilitate the relation learning in SGG. Moreover, for the long-tail problem widely studied in SGG task, it is also crucial to deal with the limited types and quantity of triplets in tail predicates. Accordingly, in this paper, we propose a Dual-granularity Relation Modeling (DRM) network to leverage fine-grained triplet cues besides the coarse-grained predicate ones. DRM utilizes contexts and semantics of predicate and triplet with Dual-granularity Constraints, generating compact and balanced representations from two perspectives to facilitate relation recognition. Furthermore, a Dual-granularity Knowledge Transfer (DKT) strategy is introduced to transfer variation from head predicates/triplets to tail ones, aiming to enrich the pattern diversity of tail classes to alleviate the long-tail problem. Extensive experiments demonstrate the effectiveness of our method, which establishes new state-of-the-art performance on Visual Genome, Open Image, and GQA datasets. Our code is available at https://github . com/jkli1998/DRM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1df4b32-543a-4067-b5e5-9c8cdea38abbCited by top-tier papers8
- Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph GenerationChangsheng Lv, Zijian Fu, Mengshi QiCVPR 2026 · 4 citations
- End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationLiwei Wang, Yanduo Zhang, Tao Lu, Fang Liu et al.ICCV 2025 · 1 citation
- Synergistic Space-Vision Processing for Predicate InferenceZhenhua Lei, Zefang Han, yu qiuICML 2026
- APT: Towards Universal Scene Graph Generation via Plug-in Adaptive Prompt TuningRuikun Luo, Changwei Gu, Jing Yang, Yuan Gao et al.ICLR 2026
- Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph GenerationTao Liu, Rongjie Li, Chongyu Wang, Xuming HeAAAI 2025
Builds on28
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu et al.CVPR 2022 · 116 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph GenerationLin Li, Long Chen, Yifeng Huang, Zhimeng Zhang et al.CVPR 2022 · 103 citations
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu et al.ICCV 2021 · 96 citations
- Context-aware Scene Graph Generation with Seq2Seq TransformersYichao Lu, Himanshu Rai, Jason Chang, Boris Knyazev et al.ICCV 2021 · 93 citations
Related papers
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 2 citations
- Memory-Based Network for Scene Graph with Unbalanced RelationsWeitao Wang, Ruyang Liu, Meng Wang, Sen Wang et al.ACM MM 2020 · 11 citations
- Iterative Learning with Extra and Inner Knowledge for Long-tail Dynamic Scene Graph GenerationYiming Li, Xiaoshan Yang, Changsheng XuACM MM 2023 · 1 citation
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 66 citations
- RelTransformer: A Transformer-Based Long-Tail Visual Relationship RecognitionJun Chen, Aniket Agarwal, Sherif Abdelkarim, Deyao Zhu et al.CVPR 2022 · 21 citations
