Vision Relation Transformer for Unbiased Scene Graph Generation
Gopika Sudhakaran, Devendra Singh Dhami, Kristian Kersting, Stefan Roth
摘要
Recent years have seen a growing interest in Scene Graph Generation (SGG), a comprehensive visual scene understanding task that aims to predict entity relationships using a relation encoder-decoder pipeline stacked on top of an object encoder-decoder backbone. Unfortunately, current SGG methods suffer from an information loss regarding the entities' local-level cues during the relation encoding process. To mitigate this, we introduce the Vision rElation TransfOrmer (VETO), consisting of a novel local-level entity relation encoder. We further observe that many existing SGG methods claim to be unbiased, but are still biased towards either head or tail classes. To overcome this bias, we introduce a Mutually Exclusive ExperT (MEET) learning strategy that captures important relation features without bias towards head or tail classes. Experimental results on the VG and GQA datasets demonstrate that VETO + MEET boosts the predictive performance by up to 47% over the state of the art while being ∼ 10× smaller. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- DeiSAM: Segment Anything with Deictic PromptingHikaru Shindo, Manuel Brack, Gopika Sudhakaran, Devendra Singh Dhami 等NeurIPS 2024 · 被引用 9 次
- Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph GenerationThong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T. Nguyen 等AAAI 2025 · 被引用 8 次
- Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term FrequencyHyeongjin Kim, Sangwon Kim, Dasom Ahn, Jong Taek Lee 等ICML 2024 · 被引用 8 次
- Interaction-Centric Knowledge Infusion and Transfer for Open Vocabulary Scene Graph GenerationLin Li, Chuhan Zhang, Dong Zhang, Chong Sun 等NeurIPS 2025 · 被引用 1 次
- EGTR: Extracting Graph from Transformer for Scene Graph GenerationJinbae Im, JeongYeon Nam, Nokyung Park, Hyungmin Lee 等CVPR 2024
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu 等CVPR 2022 · 被引用 116 次
- PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationShaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang 等ACM MM 2020 · 被引用 115 次
- Recovering the Unbiased Scene Graphs from the Biased OnesMeng-Jiun Chiou, Henghui Ding, Hanshu Yan, Changhu Wang 等ACM MM 2021 · 被引用 107 次
- Learning of Visual Relations: The Devil is in the TailsAlakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno VasconcelosICCV 2021 · 被引用 100 次
相关 Paper
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 被引用 108 次
- Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow MatchingXin Hu, Ke Qin, Wen Yin, Yuan-Fang Li 等CVPR 2026
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 被引用 2 次
- Learning to Generate an Unbiased Scene Graph by Using Attribute-Guided Predicate FeaturesLei Wang, Zejian Yuan, Badong ChenAAAI 2023 · 被引用 8 次
