Mask and Predict: Multi-step Reasoning for Scene Graph Generation
Hongshuo Tian, Ning Xu, An-An Liu, Chenggang Yan, Zhendong Mao, Quan Zhang, Yongdong Zhang
Abstract
Scene Graph Generation (SGG) aims to parse the image as a set of semantics, containing objects and their relations. Currently, the SGG methods only stay at presenting the intuitive detection in the image, such as the triplet "logo on board". Intuitively, we humans can further refine these intuitive detections as rational descriptions like "flower painted on surfboard". However, most of existing methods always formulate SGG as a straightforward task, only limited by the manner of one-time prediction, which focuses on a single-pass pipeline and predicts all the semantic. Therefore, to handle this problem, we propose a novel multi-step reasoning manner for SGG. Concretely, we break SGG into two explicit learning stages, including intuitive training stage (ITS) and rational training stage (RTS). In the first stage, we follow the traditional SGG processing to detect objects and relationships, yielding an intuitive scene graph. In the second stage, we perform multi-step reasoning to refine the intuitive scene graph. For each step of reasoning, it consists of two kinds of operations: mask and predict. According to primary predictions and their confidences, we constantly select and mask the low-confidence predictions, which features are optimized and predicted again. After several iterations, all of intuitive semantics will gradually tend to be revised with high confidences, yielding a rational scene graph. Extensive experiments on Visual Genome prove the superiority of the proposed method. Additional ablation studies and visualization cases further validate its effectiveness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f1a0f91b-6b9b-468a-9d66-66e4fef1bf38Cited by top-tier papers5
- Integrating Object-aware and Interaction-aware Knowledge for Weakly Supervised Scene Graph GenerationXingchen Li, Long Chen, Wenbo Ma, Yi Yang et al.ACM MM 2022 · 22 citations
- Improving Scene Graph Generation with Superpixel-Based Interaction LearningJingyi Wang, Can Zhang, Jinfa Huang, Botao Ren et al.ACM MM 2023 · 8 citations
- Consistent Scene Graph Generation by Constraint OptimizationBoqi Chen, Kristóf Marussy, Sebastian Pilarski, Oszkár Semeráth et al.ASE 2022 · 5 citations
- HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph GenerationCe Zhang, Simon Stepputtis, Joseph Campbell, Katia P. Sycara et al.CVPR 2024
- Leveraging Predicate and Triplet Learning for Scene Graph GenerationJiankai Li, Yunhong Wang, Xiefan Guo, Ruijie Yang et al.CVPR 2024
Related papers
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 2 citations
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 66 citations
- Not All Relations are Equal: Mining Informative Labels for Scene Graph GenerationArushi Goel, Basura Fernando, Frank Keller, Hakan BilenCVPR 2022 · 30 citations
- HL-Net: Heterophily Learning Network for Scene Graph GenerationXin Lin, Changxing Ding, Yibing Zhan, Zijian Li et al.CVPR 2022 · 51 citations
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 47 citations
