Iterative Scene Graph Generation
Siddhesh Khandelwal, Leonid Sigal
摘要
The task of scene graph generation entails identifying object entities and their corresponding interaction predicates in a given image (or video). Due to the combinatorially large solution space, existing approaches to scene graph generation assume certain factorization of the joint distribution to make the estimation feasible (e.g., assuming that objects are conditionally independent of predicate predictions). However, this fixed factorization is not ideal under all scenarios (e.g., for images where an object entailed in interaction is small and not discernible on its own). In this work, we propose a novel framework for scene graph generation that addresses this limitation, as well as introduces dynamic conditioning on the image, using message passing in a Markov Random Field. This is implemented as an iterative refinement procedure wherein each modification is conditioned on the graph generated in the previous iteration. This conditioning across refinement steps allows joint reasoning over entities and relations. This framework is realized via a novel and end-to-end trainable transformer-based architecture. In addition, the proposed framework can improve existing approach performance. Through extensive experiments on Visual Genome [29] and Action Genome [24] benchmark datasets we show improved performance on the scene graph generation task. Contributions. To realize the aforementioned iterative framework, we propose a novel and intuitive transformer [46] based architecture. On a technical level, our model defines three separate multi-layer multi-output synchronized decoders, wherein each decoder layer is tasked with modeling either the subject, object, or predicate components of a relationship triplets. Therefore, the combined outputs from each layer of the three decoders generates a scene graph estimate. The inputs to each decoder layer are conditioned to enable joint reasoning across decoders and effective refinement of previous layer estimates. This conditioning is achieved implicitly via a novel joint loss, and explicitly via crossdecoder and layer-wise attention. Additionally, each decoder layer is also conditioned on the image features, which are provided by a shared encoder. As our proposed model is end-to-end trainable, it addresses the limitation of two-stage approaches, allowing image features to directly adopt to the scene graph generation task. Finally, to tackle the long-tail nature of the scene graph predicate classes [10], we employ a loss weighting strategy to enable flexible trade-off between dominant (head) and underrepresented (tail) predicate classes in the long-tail distribution. In contrast to data sampling strategies [10, 31] , this has a benefit of not requiring additional fine-tuning of models with sampled data post training. We illustrate that our proposed architecture achieves state-of-the-art performance on two benchmark datasets -Visual Genome [29] and Action Genome [24] ; and thoroughly analyze effectiveness of the approach as a function of the refinement steps, design choices employed and as a generic add-on to an existing, MOTIF [53], architecture. Related Work Scene Graph Generation. Scene graph generation has emerged as a popular research area in the vision community [10, 27, 30, 34, 38, 39, 44, 45, 48, 50, 52, 53] . Existing scene graph generation methods can be broadly categorized as either one-stage or two-stage approaches. The first step of the predominant approach -the two-stage methods -involves pre-training a strong object detector for all object classes in the dataset, usually using detector like Faster-RCNN [40] . The graph generation network is then built on top of the object information (bounding boxes and corresponding
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren 等NeurIPS 2024 · 被引用 13 次
- Focusing on Flexible Masks: A Novel Framework for Panoptic Scene Graph Generation with Relation ConstraintsJiarui Yang, Chuan Wang, Zeming Liu, Jiahong Wu 等ACM MM 2023 · 被引用 8 次
- GraphMorph: Tubular Structure Extraction by Morphing Predicted GraphsZhao Zhang, Ziwei Zhao, Dong Wang, Liwei WangNeurIPS 2024 · 被引用 4 次
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 被引用 2 次
- IS-GGT: Iterative Scene Graph Generation with Generative TransformersSanjoy Kundu, Sathyanarayanan N. AakurCVPR 2023
它引用的顶会 Paper20
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao 等ICCV 2019 · 被引用 191 次
相关 Paper
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 被引用 108 次
- OED: Towards One-stage End-to-End Dynamic Scene Graph GenerationGuan Wang, Zhimin Li, Qingchao Chen, Yang LiuCVPR 2024 · 被引用 12 次
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn 等ICCV 2021 · 被引用 163 次
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 被引用 38 次
- Context-aware Scene Graph Generation with Seq2Seq TransformersYichao Lu, Himanshu Rai, Jason Chang, Boris Knyazev 等ICCV 2021 · 被引用 93 次
