HSGG: Training-Free Hierarchical Scene Graph Generation with Geometry-Guided Relation Reasoning
yunzhe Liu, Wenbiao Liu, Lihui Cen, Zhe Qu, Yigang Cen, Yixiong Liang, Yongfang Xie, Shichao Kan
摘要
Scene Graph Generation (SGG) connects visual perception with structured reasoning, but is limited by scarce annotations and the long-tailed distribution of relational predicates. Training-free methods based on vision-language models (VLMs) reduce supervision requirements, yet often rely on flat graph structures and produce hallucinated relations. We present HSGG, a training-free framework for open-world Hierarchical Scene Graph Generation, whose inference proceeds in two steps. First, bidirectional hierarchical entity perception combines top-down object expansion with bottom-up attribute reasoning to construct multi-level scene hierarchies capturing part–whole semantics. Second, geometry-guided relation reasoning infers valid relations from these structured entities: geometry-aware relation filtering first prunes spatially implausible object pairs using 2D proximity, depth cues, and object scale, and geometry-grounded contrastive relation decoding then suppresses hallucinated predicates by contrasting predictions from a visually grounded expert against a hallucination-prone geometric prior, ensuring relations are both geometrically consistent and semantically coherent. Experiments show that HSGG generalizes effectively to unseen objects and predicates without training, substantially reduces relational hallucinations, and consistently improves downstream reasoning performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen 等ICLR 2026 · 被引用 720 次
- Analyzing and Mitigating Object Hallucination in Large Vision-Language ModelsYiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang 等ICLR 2024 · 被引用 316 次
相关 Paper
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen 等CVPR 2024
- Seeing Beyond Hallucinations: LLM-based Compositional Information Extraction for Multimodal ReasoningXinwei Li, Li Lin, Shuai Wang, Hanqian WuSIGIR 2025 · 被引用 1 次
- Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningFei Yu, Quan Deng, Shengeng Tang, Yuehua Li 等AAAI 2026 · 被引用 2 次
- ReLaGS: Relational Language Gaussian SplattingYaxu Xie, Abdalla Arafa, Alireza Javanmardi, Christen Millerdurai 等CVPR 2026 · 被引用 7 次
- Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language CompositionalityHarman Singh, Pengchuan Zhang, Qifan Wang, Mengjiao Wang 等EMNLP 2023 · 被引用 10 次
