Hybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph Generation
Jiawei Fu, Tiantian Zhang, Kai Chen, Qi Dou
摘要
Scene graph generation is a pivotal task in computer vision, focusing on comprehensive identification of visual relation tuples embedded within images. The advancement of methods involving triplets has sought to enhance task performance by integrating triplets as contextual features for more precise predicate identification from component level. However, challenges remain due to interference from multirole objects in overlapping tuples within complex environments, which impairs the model's ability to distinguish and align specific triplet features for reasoning diverse semantics of multi-role objects. To address these issues, we introduce a novel framework that incorporates a triplet alignment model into a hybrid reciprocal transformer architecture, starting from using triplet mask features to guide the learning of component-level relation graphs. To effectively distinguish multi-role objects characterized by overlapping visual relation tuples, we introduce a triplet alignment loss, which provides multi-role objects with aligned features from triplet and helps customize them. Additionally, we explore the inherent connectivity between hybrid aligned triplet and component features through a bidirectional refinement module, which enhances feature interaction and reciprocal reinforcement. Experimental results demonstrate that our model achieves state-of-the-art performance on the Visual Genome and Action Genome datasets, underscoring its effectiveness and adaptability. Project page: hq-sg.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow MatchingXin Hu, Ke Qin, Wen Yin, Yuan-Fang Li 等CVPR 2026
- APT: Towards Universal Scene Graph Generation via Plug-in Adaptive Prompt TuningRuikun Luo, Changwei Gu, Jing Yang, Yuan Gao 等ICLR 2026
- Learning Gaussian Mixture-distributed Prototypes for 3D Scene Graph Generation from RGB-D SequencesRongxing Ding, Hongyu Qu, Xinguang Xiang, Pengpeng Li 等ICML 2026
它引用的顶会 Paper35
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo 等CVPR 2022 · 被引用 879 次
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 被引用 108 次
- The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph GenerationLin Li, Long Chen, Yifeng Huang, Zhimeng Zhang 等CVPR 2022 · 被引用 103 次
- Learning of Visual Relations: The Devil is in the TailsAlakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno VasconcelosICCV 2021 · 被引用 100 次
相关 Paper
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 被引用 2 次
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 被引用 47 次
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 被引用 66 次
- Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph GenerationYiming Li, Sisi You, Bing-Kun BaoCVPR 2026
- IS-GGT: Iterative Scene Graph Generation with Generative TransformersSanjoy Kundu, Sathyanarayanan N. AakurCVPR 2023
