Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph Generation
Xingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu, Yuan Cheng, Liqiang Nie
Abstract
Scene Graph Generation, which generally follows a regular encoder-decoder pipeline, aims to first encode the visual contents within the given image and then parse them into a compact summary graph. Existing SGG approaches generally not only neglect the insufficient modality fusion between vision and language, but also fail to provide informative predicates due to the biased relationship predictions, leading SGG far from practical. Towards this end, we first present a novel Stacked Hybrid-Attention network, which facilitates the intra-modal refinement as well as the intermodal interaction, to serve as the encoder. We then devise an innovative Group Collaborative Learning strategy to optimize the decoder. Particularly, based on the observation that the recognition capability of one classifier is limited towards an extremely unbalanced dataset, we first deploy a group of classifiers that are expert in distinguishing different subsets of classes, and then cooperatively optimize them from two aspects to promote the unbiased SGG. Experiments conducted on VG and GQA datasets demonstrate that, we not only establish a new state-of-the-art in the unbiased metric, but also nearly double the performance compared with two baselines. Our code is available at https://github.com/dongxingning/SHA-GCL-for-SGG .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 372c84da-145d-480b-9568-2d38ce2f6107Cited by top-tier papers33
- Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsLin Li, Jun Xiao, Guikun Chen, Jian Shao et al.NeurIPS 2023 · 52 citations
- Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open WorldQifan Yu, Juncheng Li, Yu Wu, Siliang Tang et al.ICCV 2023 · 51 citations
- Unbiased Heterogeneous Scene Graph Generation with Relation-Aware Message Passing Neural NetworkKanghoon Yoon, Kibum Kim, Jinyoung Moon, Chanyoung ParkAAAI 2023 · 48 citations
- Compositional Feature Augmentation for Unbiased Scene Graph GenerationLin Li, Guikun Chen, Jun Xiao, Yi Yang et al.ICCV 2023 · 36 citations
- HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph GenerationZijian Zhou, Miaojing Shi, Holger CaesarICCV 2023 · 29 citations
Builds on18
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationShaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang et al.ACM MM 2020 · 115 citations
- Recovering the Unbiased Scene Graphs from the Biased OnesMeng-Jiun Chiou, Henghui Ding, Hanshu Yan, Changhu Wang et al.ACM MM 2021 · 107 citations
- Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical StudyZhiqiang Shen, Zechun Liu, Dejia Xu, Zitian Chen et al.ICLR 2021 · 83 citations
Related papers
- Learning to Generate an Unbiased Scene Graph by Using Attribute-Guided Predicate FeaturesLei Wang, Zejian Yuan, Badong ChenAAAI 2023 · 8 citations
- Vision Relation Transformer for Unbiased Scene Graph GenerationGopika Sudhakaran, Devendra Singh Dhami, Kristian Kersting, Stefan RothICCV 2023 · 27 citations
- Synergetic Prototype Learning Network for Unbiased Scene Graph GenerationRuonan Zhang, Ziwei Shang, Fengjuan Wang, Zhaoqilin Yang et al.ACM MM 2024 · 5 citations
- Predicate Debiasing in Vision-Language Models Integration for Scene Graph Generation EnhancementYuxuan Wang, Xiaoyuan LiuEMNLP 2024 · 1 citation
- Weakly-supervised Video Scene Graph Generation via Unbiased Cross-modal LearningZiyue Wu, Junyu Gao, Changsheng XuACM MM 2023 · 5 citations
