Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph Generation
Changsheng Lv, Zijian Fu, Mengshi Qi
Abstract
In this paper, we propose Robo-SGG, a plug-and-play module for robust scene graph generation (SGG). Unlike standard SGG, the robust scene graph generation aims to perform inference on a diverse range of corrupted images, with the core challenge being the domain shift between the clean and corrupted images. Existing SGG methods suffer from degraded performance due to shifted visual features (e.g., corruption interference or occlusions). To obtain robust visual features, we leverage layout information, representing the global structure of an image, which is robust to domain shift, to enhance the robustness of SGG methods under corruption. Specifically, we employ Instance Normalization (IN) to alleviate the domain-specific variations and recover the robust structural features (i.e., the positional and semantic relationships among objects) by the proposed Layout-Oriented Restitution. Furthermore, under corrupted images, we introduce a Layout-Embedded Encoder (LEE) that adaptively fuses layout and visual features via a gating mechanism, enhancing the robustness of positional and semantic representations for objects and predicates. Note that our proposed Robo-SGG module is designed as a plug-and-play component, which can be easily integrated into any baseline SGG model. Extensive experiments demonstrate that by integrating the state-of-the-art method into our proposed Robo-SGG, we achieve relative improvements of 6.3%, 11.1%, and 8.0% in mR@50 for PredCls, SGCls, and SGDet tasks on the VG-C benchmark, respectively, and achieve new state-of-the-art performance in the corruption scene graph generation benchmark (VG-C and GQA-C). We will release our source code and model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2993480c-9f9d-4cdb-9a93-0d7c40385769Cited by top-tier papers1
Ask how each one uses itBuilds on30
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 997 citations
- On Interaction Between Augmentations and Corruptions in Natural Corruption RobustnessEric Mintun, Alexander Kirillov, Saining XieNeurIPS 2021 · 138 citations
- The Norm Must Go On: Dynamic Unsupervised Domain Adaptation by NormalizationMuhammad Jehanzeb Mirza, Jakub Micorek, Horst Possegger, Horst BischofCVPR 2022 · 119 citations
Related papers
- HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph GenerationCe Zhang, Simon Stepputtis, Joseph Campbell, Katia P. Sycara et al.CVPR 2024
- Scene Graph-Grounded Image GenerationFuyun Wang, Tong Zhang, Yuanzhi Wang, Xiaoya Zhang et al.AAAI 2025 · 1 citation
- Heterogeneous Learning for Scene Graph GenerationYunqing He, Tongwei Ren, Jinhui Tang, Gangshan WuACM MM 2022 · 3 citations
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu et al.ICCV 2021 · 96 citations
- Object-Centric Image Generation from LayoutsTristan Sylvain, Pengchuan Zhang, Yoshua Bengio, R. Devon Hjelm et al.AAAI 2021 · 107 citations
