Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph Generation
Xingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu, Yuan Cheng, Liqiang Nie
摘要
Scene Graph Generation, which generally follows a regular encoder-decoder pipeline, aims to first encode the visual contents within the given image and then parse them into a compact summary graph. Existing SGG approaches generally not only neglect the insufficient modality fusion between vision and language, but also fail to provide informative predicates due to the biased relationship predictions, leading SGG far from practical. Towards this end, we first present a novel Stacked Hybrid-Attention network, which facilitates the intra-modal refinement as well as the intermodal interaction, to serve as the encoder. We then devise an innovative Group Collaborative Learning strategy to optimize the decoder. Particularly, based on the observation that the recognition capability of one classifier is limited towards an extremely unbalanced dataset, we first deploy a group of classifiers that are expert in distinguishing different subsets of classes, and then cooperatively optimize them from two aspects to promote the unbiased SGG. Experiments conducted on VG and GQA datasets demonstrate that, we not only establish a new state-of-the-art in the unbiased metric, but also nearly double the performance compared with two baselines. Our code is available at https://github.com/dongxingning/SHA-GCL-for-SGG .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsLin Li, Jun Xiao, Guikun Chen, Jian Shao 等NeurIPS 2023 · 被引用 52 次
- Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open WorldQifan Yu, Juncheng Li, Yu Wu, Siliang Tang 等ICCV 2023 · 被引用 51 次
- Unbiased Heterogeneous Scene Graph Generation with Relation-Aware Message Passing Neural NetworkKanghoon Yoon, Kibum Kim, Jinyoung Moon, Chanyoung ParkAAAI 2023 · 被引用 48 次
- Compositional Feature Augmentation for Unbiased Scene Graph GenerationLin Li, Guikun Chen, Jun Xiao, Yi Yang 等ICCV 2023 · 被引用 36 次
- HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph GenerationZijian Zhou, Miaojing Shi, Holger CaesarICCV 2023 · 被引用 29 次
它引用的顶会 Paper18
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao 等ICCV 2019 · 被引用 191 次
- PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationShaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang 等ACM MM 2020 · 被引用 115 次
- Recovering the Unbiased Scene Graphs from the Biased OnesMeng-Jiun Chiou, Henghui Ding, Hanshu Yan, Changhu Wang 等ACM MM 2021 · 被引用 107 次
- Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical StudyZhiqiang Shen, Zechun Liu, Dejia Xu, Zitian Chen 等ICLR 2021 · 被引用 83 次
相关 Paper
- Learning to Generate an Unbiased Scene Graph by Using Attribute-Guided Predicate FeaturesLei Wang, Zejian Yuan, Badong ChenAAAI 2023 · 被引用 8 次
- Vision Relation Transformer for Unbiased Scene Graph GenerationGopika Sudhakaran, Devendra Singh Dhami, Kristian Kersting, Stefan RothICCV 2023 · 被引用 27 次
- Synergetic Prototype Learning Network for Unbiased Scene Graph GenerationRuonan Zhang, Ziwei Shang, Fengjuan Wang, Zhaoqilin Yang 等ACM MM 2024 · 被引用 5 次
- Predicate Debiasing in Vision-Language Models Integration for Scene Graph Generation EnhancementYuxuan Wang, Xiaoyuan LiuEMNLP 2024 · 被引用 1 次
- Weakly-supervised Video Scene Graph Generation via Unbiased Cross-modal LearningZiyue Wu, Junyu Gao, Changsheng XuACM MM 2023 · 被引用 5 次
