SGTR: End-to-end Scene Graph Generation with Transformer
Rongjie Li, Songyang Zhang, Xuming He
摘要
Scene Graph Generation (SGG) remains a challenging visual understanding task due to its compositional property. Most previous works adopt a bottom-up two-stage or a point-based one-stage approach, which often suffers from high time complexity or sub-optimal designs. In this work, we propose a novel SGG method to address the aforementioned issues, formulating the task as a bipartite graph construction problem. To solve the problem, we develop a transformer-based end-to-end framework that first generates the entity and predicate proposal set, followed by inferring directed edges to form the relation triplets. In particular, we develop a new entity-aware predicate representation based on a structural predicate generator that leverages the compositional property of relationships. Moreover, we design a graph assembling module to infer the connectivity of the bipartite scene graph based on our entity-aware structure, enabling us to generate the scene graph in an end-to-end manner. Extensive experimental results show that our design is able to achieve the state-of-the-art or comparable performance on two challenging benchmarks, surpassing most of the existing approaches and enjoying higher efficiency in inference. We hope our model can serve as a strong baseline for the Transformer-based scene graph generation. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Code is available: https://github.com/Scarecrow0/SGTR
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper50
- RLIPv2: Fast Scaling of Relational Language-Image Pre-trainingHangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Albanie 等ICCV 2023 · 被引用 69 次
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 被引用 47 次
- Scene Graph Generation with Role-Playing Large Language ModelsGuikun Chen, Jin Li, Wenguan WangNeurIPS 2024 · 被引用 33 次
- 4D Panoptic Scene Graph GenerationJingkang Yang, Jun Cen, Wenxuan Peng, Shuai Liu 等NeurIPS 2023 · 被引用 33 次
- HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph GenerationZijian Zhou, Miaojing Shi, Holger CaesarICCV 2023 · 被引用 29 次
它引用的顶会 Paper26
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Mining the Benefits of Two-stage and One-stage HOI DetectionAixi Zhang, Yue Liao, Si Liu, Miao Lu 等NeurIPS 2021 · 被引用 218 次
- PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationShaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang 等ACM MM 2020 · 被引用 115 次
- Recovering the Unbiased Scene Graphs from the Biased OnesMeng-Jiun Chiou, Henghui Ding, Hanshu Yan, Changhu Wang 等ACM MM 2021 · 被引用 107 次
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu 等ICCV 2021 · 被引用 96 次
相关 Paper
- IS-GGT: Iterative Scene Graph Generation with Generative TransformersSanjoy Kundu, Sathyanarayanan N. AakurCVPR 2023
- Weakly Supervised Visual Semantic ParsingAlireza Zareian, Svebor Karaman, Shih-Fu ChangCVPR 2020
- Single-Stage Visual Relationship Learning using Conditional QueriesAlakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno VasconcelosNeurIPS 2022 · 被引用 11 次
- SGFormer: Semantic Graph Transformer for Point Cloud-Based 3D Scene Graph GenerationChangsheng Lv, Mengshi Qi, Xia Li, Zhengyuan Yang 等AAAI 2024 · 被引用 32 次
- DSGG: Dense Relation Transformer for an End-to-End Scene Graph GenerationZeeshan Hayder, Xuming HeCVPR 2024
