HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph Generation
Zijian Zhou, Miaojing Shi, Holger Caesar
摘要
Panoptic Scene Graph generation (PSG) is a recently proposed task in image scene understanding that aims to segment the image and extract triplets of subjects, objects and their relations to build a scene graph. This task is particularly challenging for two reasons. First, it suffers from a long-tail problem in its relation categories, making naive biased methods more inclined to high-frequency relations. Existing unbiased methods tackle the long-tail problem by data/loss rebalancing to favor low-frequency relations. Second, a subject-object pair can have two or more semantically overlapping relations. While existing methods favor one over the other, our proposed HiLo framework lets different network branches specialize on low and high frequency relations, enforce their consistency and fuse the results. To the best of our knowledge we are the first to propose an explicitly unbiased PSG method. In extensive experiments we show that our HiLo framework achieves state-of-the-art results on the PSG task. We also apply our method to the Scene Graph Generation task that predicts boxes instead of masks and see improvements over all baseline methods. Code is available at https://github. com/franciszzj/HiLo .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationLinfeng Yuan, Miaojing Shi, Zijie Yue, Qijun ChenCVPR 2024 · 被引用 12 次
- SPADE: Spatial-Aware Denoising Network for Open-Vocabulary Panoptic Scene Graph Generation with Long- and Local-Range Context ReasoningXin Hu, Ke Qin, Guiduo Duan, Ming Li 等ICCV 2025 · 被引用 1 次
- Universal Scene Graph GenerationShengqiong Wu, Hao Fei, Tat-Seng ChuaCVPR 2025
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu 等CVPR 2022 · 被引用 116 次
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 被引用 108 次
相关 Paper
- Focusing on Flexible Masks: A Novel Framework for Panoptic Scene Graph Generation with Relation ConstraintsJiarui Yang, Chuan Wang, Zeming Liu, Jiahong Wu 等ACM MM 2023 · 被引用 8 次
- Biased-Predicate Annotation Identification via Unbiased Visual Predicate RepresentationLi Li, Chenwei Wang, You Qin, Wei Ji 等ACM MM 2023 · 被引用 13 次
- SegPVSG: Panoptic Video Scene Graph Generation via Temporal Focusing and Generative AugmentationYiKai Li, Quhui Ke, Jinglin Liang, Zhiyuan Zhang 等ICML 2026
- TextPSG: Panoptic Scene Graph Generation from Textual DescriptionsChengyang Zhao, Yikang Shen, Zhenfang Chen, Mingyu Ding 等ICCV 2023 · 被引用 24 次
- DSGG: Dense Relation Transformer for an End-to-End Scene Graph GenerationZeeshan Hayder, Xuming HeCVPR 2024
