Vision Relation Transformer for Unbiased Scene Graph Generation
Gopika Sudhakaran, Devendra Singh Dhami, Kristian Kersting, Stefan Roth
Abstract
Recent years have seen a growing interest in Scene Graph Generation (SGG), a comprehensive visual scene understanding task that aims to predict entity relationships using a relation encoder-decoder pipeline stacked on top of an object encoder-decoder backbone. Unfortunately, current SGG methods suffer from an information loss regarding the entities' local-level cues during the relation encoding process. To mitigate this, we introduce the Vision rElation TransfOrmer (VETO), consisting of a novel local-level entity relation encoder. We further observe that many existing SGG methods claim to be unbiased, but are still biased towards either head or tail classes. To overcome this bias, we introduce a Mutually Exclusive ExperT (MEET) learning strategy that captures important relation features without bias towards head or tail classes. Experimental results on the VG and GQA datasets demonstrate that VETO + MEET boosts the predictive performance by up to 47% over the state of the art while being ∼ 10× smaller. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bef92f2-8f3b-4d03-bf04-4f3baa5ca73fCited by top-tier papers13
- DeiSAM: Segment Anything with Deictic PromptingHikaru Shindo, Manuel Brack, Gopika Sudhakaran, Devendra Singh Dhami et al.NeurIPS 2024 · 9 citations
- Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph GenerationThong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T. Nguyen et al.AAAI 2025 · 8 citations
- Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term FrequencyHyeongjin Kim, Sangwon Kim, Dasom Ahn, Jong Taek Lee et al.ICML 2024 · 8 citations
- Interaction-Centric Knowledge Infusion and Transfer for Open Vocabulary Scene Graph GenerationLin Li, Chuhan Zhang, Dong Zhang, Chong Sun et al.NeurIPS 2025 · 1 citation
- EGTR: Extracting Graph from Transformer for Scene Graph GenerationJinbae Im, JeongYeon Nam, Nokyung Park, Hyungmin Lee et al.CVPR 2024
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu et al.CVPR 2022 · 116 citations
- PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationShaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang et al.ACM MM 2020 · 115 citations
- Recovering the Unbiased Scene Graphs from the Biased OnesMeng-Jiun Chiou, Henghui Ding, Hanshu Yan, Changhu Wang et al.ACM MM 2021 · 107 citations
- Learning of Visual Relations: The Devil is in the TailsAlakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno VasconcelosICCV 2021 · 100 citations
Related papers
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow MatchingXin Hu, Ke Qin, Wen Yin, Yuan-Fang Li et al.CVPR 2026
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 2 citations
- Learning to Generate an Unbiased Scene Graph by Using Attribute-Guided Predicate FeaturesLei Wang, Zejian Yuan, Badong ChenAAAI 2023 · 8 citations
