Single-Stage Visual Relationship Learning using Conditional Queries
Alakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno Vasconcelos
Abstract
Research in scene graph generation (SGG) usually considers two-stage models, that is, detecting a set of entities, followed by combining them and labeling all possible relationships. While showing promising results, the pipeline structure induces large parameter and computation overhead, and typically hinders end-to-end optimizations. To address this, recent research attempts to train single-stage models that are computationally efficient. With the advent of DETR[3], a set based detection model, one-stage models attempt to predict a set of subject-predicate-object triplets directly in a single shot. However, SGG is inherently a multi-task learning problem that requires modeling entity and predicate distributions simultaneously. In this paper, we propose Transformers with conditional queries for SGG, namely, TraCQ with a new formulation for SGG that avoids the multi-task learning problem and the combinatorial entity pair distribution. We employ a DETR-based encoder-decoder design and leverage conditional queries to significantly reduce the entity label space as well, which leads to 20% fewer parameters compared to state-of-the-art singlestage models. Experimental results show that TraCQ not only outperforms existing single-stage scene graph generation methods, it also beats many state-of-the-art two-stage methods on the Visual Genome dataset, yet is capable of end-to-end training and faster inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de0156d5-2fef-4302-b234-279adfe36b4cCited by top-tier papers3
- Transitivity Recovering Decompositions: Interpretable and Robust Fine-Grained RelationshipsAbhra Chaudhuri, Massimiliano Mancini, Zeynep Akata, Anjan DuttaNeurIPS 2023 · 5 citations
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 2 citations
- DSGG: Dense Relation Transformer for an End-to-End Scene Graph GenerationZeeshan Hayder, Xuming HeCVPR 2024
Builds on12
- PnP-DETR: Towards Efficient Visual Analysis with TransformersTao Wang, Li Yuan, Yunpeng Chen, Jiashi Feng et al.ICCV 2021 · 125 citations
- PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationShaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang et al.ACM MM 2020 · 115 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- Learning of Visual Relations: The Devil is in the TailsAlakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno VasconcelosICCV 2021 · 100 citations
- Context-aware Scene Graph Generation with Seq2Seq TransformersYichao Lu, Himanshu Rai, Jason Chang, Boris Knyazev et al.ICCV 2021 · 93 citations
Related papers
- EGTR: Extracting Graph from Transformer for Scene Graph GenerationJinbae Im, JeongYeon Nam, Nokyung Park, Hyungmin Lee et al.CVPR 2024
- Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph GenerationMinghan Chen, Guikun Chen, Wenguan Wang, Yi YangICLR 2025
- IS-GGT: Iterative Scene Graph Generation with Generative TransformersSanjoy Kundu, Sathyanarayanan N. AakurCVPR 2023
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 47 citations
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 66 citations
