Synergistic Space-Vision Processing for Predicate Inference
Zhenhua Lei, Zefang Han, yu qiu
摘要
Scene graph generation (SGG) aims to parse an image into a structured graph of objects and their predicates, enabling explicit relational reasoning for visual understanding. However, prevailing methods often over-predict geometric predicates, resulting in scene graphs that are factually correct yet semantically shallow. While recent works effectively attribute this phenomenon to the long-tailed data distribution, we identify another critical factor driving such biased prediction: co-occurrence-induced representation entanglement, where geometric and non-geometric predicates that frequently co-occur are encoded into overly similar representations. To this end, we introduce Dual-stream Synergistic Network (DS-Net) that models geometric and non-geometric predicates with two specialized streams, coupled with a bidirectional cross-stream fusion mechanism. The space stream focuses on spatial and structural cues, while the vision stream captures fine-grained visual evidence and semantic priors. Extensive experiments show that DS-Net consistently improves predicate inference, achieving 1.3% 6.1% absolute gains in mR@100 on the SGGen task when integrated into existing SGG methods. These results highlight the importance of synergistic modeling of geometric and non-geometric predicates for generating semantically richer scene graphs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving ScenarioTianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao 等AAAI 2024 · 被引用 314 次
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu 等CVPR 2022 · 被引用 116 次
- Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question AnsweringWeizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca 等NeurIPS 2023 · 被引用 108 次
- The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph GenerationLin Li, Long Chen, Yifeng Huang, Zhimeng Zhang 等CVPR 2022 · 被引用 103 次
- Context-aware Scene Graph Generation with Seq2Seq TransformersYichao Lu, Himanshu Rai, Jason Chang, Boris Knyazev 等ICCV 2021 · 被引用 93 次
相关 Paper
- Leveraging Predicate and Triplet Learning for Scene Graph GenerationJiankai Li, Yunhong Wang, Xiefan Guo, Ruijie Yang 等CVPR 2024
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 被引用 2 次
- Learning to Generate an Unbiased Scene Graph by Using Attribute-Guided Predicate FeaturesLei Wang, Zejian Yuan, Badong ChenAAAI 2023 · 被引用 8 次
- Weakly Supervised Visual Semantic ParsingAlireza Zareian, Svebor Karaman, Shih-Fu ChangCVPR 2020
- End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationLiwei Wang, Yanduo Zhang, Tao Lu, Fang Liu 等ICCV 2025 · 被引用 1 次
