Navigating the Unseen: Zero-shot Scene Graph Generation via Capsule-Based Equivariant Features
Wenhuan Huang, Yi Ji, Guiqian Zhu, Li Ying, Chunping Liu
Abstract
In scene graph generation (SGG), the accurate prediction of unseen triplets is essential for its effectiveness in downstream vision-language tasks. We hypothesize that the predicates of unseen triplets can be viewed as transformations of seen predicates in feature space, and the essence of the zero-shot task is to bridge the gap caused by this transformation. Traditional models, however, have difficulty addressing this challenge, which we attribute to their inability to model the predicates equivariant. To overcome this limitation, we introduce a novel framework based on capsule networks (CAPSGG). We propose a Three-Stream Pipeline that generates modality-specific representations for predicates, while building low-level predicate capsules of these modalities. Then, these capsules are aggregated into highlevel predicate capsules using a Routing Capsule Layer. In addition, we introduce GroupLoss to aggregate capsules with the same predicate label into groups. This replaces the global loss with the intra-group loss, effectively balancing the learning of predicate invariant and equivariant features while mitigating the impact of the severe long-tail distribution of the predicate categories. Our extensive experiments demonstrate the notable superiority of our approach over state-of-the-art methods, with zero-shot indicators outperforming up to 132.26% on the SGCls task than the T-CAR [21] . Our code will be available upon publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow MatchingXin Hu, Ke Qin, Wen Yin, Yuan-Fang Li et al.CVPR 2026
- HSGG: Training-Free Hierarchical Scene Graph Generation with Geometry-Guided Relation Reasoningyunzhe Liu, Wenbiao Liu, Lihui Cen, Zhe Qu et al.ICML 2026
- Learning Gaussian Mixture-distributed Prototypes for 3D Scene Graph Generation from RGB-D SequencesRongxing Ding, Hongyu Qu, Xinguang Xiang, Pengpeng Li et al.ICML 2026
Builds on12
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu et al.CVPR 2022 · 116 citations
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 66 citations
- Not All Relations are Equal: Mining Informative Labels for Scene Graph GenerationArushi Goel, Basura Fernando, Frank Keller, Hakan BilenCVPR 2022 · 30 citations
- Generative Compositional Augmentations for Scene Graph PredictionBoris Knyazev, Harm de Vries, Catalina Cangea, Graham W. Taylor et al.ICCV 2021 · 30 citations
- Environment-Invariant Curriculum Relation Learning for Fine-Grained Scene Graph GenerationYukuan Min, Aming Wu, Cheng DengICCV 2023 · 16 citations
Related papers
- Multi-view Invariance Learning for 3D Scene Graph Pre-training via Collaborative Cross-Modal RegularizationYucheng Huang, Luping Ji, Ruijie Xiao, Jiayuan SunAAAI 2026
- Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open WorldQifan Yu, Juncheng Li, Yu Wu, Siliang Tang et al.ICCV 2023 · 51 citations
- CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive LearningLianggangxu Chen, Xuejiao Wang, Jiale Lu, Shaohui Lin et al.CVPR 2024
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu et al.ICCV 2021 · 96 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
