HOSE-Net: Higher Order Structure Embedded Network for Scene Graph Generation
Meng Wei, Chun Yuan, Xiaoyu Yue, Kuo Zhong
摘要
Scene graph generation aims to produce structured representations for images, which requires to understand the relations between objects. Due to the continuous nature of deep neural networks, the prediction of scene graphs is divided into object detection and relation classification. However, the independent relation classes cannot separate the visual features well. Although some methods organize the visual features into graph structures and use message passing to learn contextual information, they still suffer from drastic intra-class variations and unbalanced data distributions. One important factor is that they learn an unstructured output space that ignores the inherent structures of scene graphs. Accordingly, in this paper, we propose a Higher Order Structure Embedded Network (HOSE-Net) to mitigate this issue. First, we propose a novel structure-aware embedding-to-classifier(SEC) module to incorporate both local and global structural information of relationships into the output space. Specifically, a set of context embeddings are learned via local graph based message passing and then mapped to a global structure based classification space. Second, since learning too many context-specific classification subspaces can suffer from data sparsity issues, we propose a hierarchical semantic aggregation(HSA) module to reduces the number of subspaces by introducing higher order structural information. HSA is also a fast and flexible tool to automatically search a semantic object hierarchy based on relational knowledge graphs. Extensive experiments show that the proposed HOSE-Net achieves the state-of-the-art performance on two popular benchmarks of Visual Genome and VRD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Recovering the Unbiased Scene Graphs from the Biased OnesMeng-Jiun Chiou, Henghui Ding, Hanshu Yan, Changhu Wang 等ACM MM 2021 · 被引用 107 次
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura 等ACM MM 2022 · 被引用 40 次
- Rethinking the Two-Stage Framework for Grounded Situation RecognitionMeng Wei, Long Chen, Wei Ji, Xiaoyu Yue 等AAAI 2022 · 被引用 38 次
- Compositional Feature Augmentation for Unbiased Scene Graph GenerationLin Li, Guikun Chen, Jun Xiao, Yi Yang 等ICCV 2023 · 被引用 36 次
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 被引用 2 次
相关 Paper
- HL-Net: Heterophily Learning Network for Scene Graph GenerationXin Lin, Changxing Ding, Yibing Zhan, Zijian Li 等CVPR 2022 · 被引用 51 次
- Target Adaptive Context Aggregation for Video Scene Graph GenerationYao Teng, Limin Wang, Zhifeng Li, Gangshan WuICCV 2021 · 被引用 80 次
- Unbiased Heterogeneous Scene Graph Generation with Relation-Aware Message Passing Neural NetworkKanghoon Yoon, Kibum Kim, Jinyoung Moon, Chanyoung ParkAAAI 2023 · 被引用 48 次
- Prototype-Based Embedding Network for Scene Graph GenerationChaofan Zheng, Xinyu Lyu, Lianli Gao, Bo Dai 等CVPR 2023
- Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic SpaceYong Zhang, Yingwei Pan, Ting Yao, Rui Huang 等CVPR 2023
