Deep Grouping Model for Unified Perceptual Parsing
Zhiheng Li, Wenxuan Bao, Jiayang Zheng, Chenliang Xu
摘要
The perceptual-based grouping process produces a hierarchical and compositional image representation that helps both human and machine vision systems recognize heterogeneous visual concepts. Examples can be found in the classical hierarchical superpixel segmentation or image parsing works. However, the grouping process is largely overlooked in modern CNN-based image segmentation networks due to many challenges, including the inherent incompatibility between the grid-shaped CNN feature map and the irregularshaped perceptual grouping hierarchy. Overcoming these challenges, we propose a deep grouping model (DGM) that tightly marries the two types of representations and defines a bottom-up and a top-down process for feature exchanging. When evaluating the model on the recent Broden+ dataset for the unified perceptual parsing task, it achieves state-ofthe-art results while having a small computational overhead compared to other contextual-based segmentation models. Furthermore, the DGM has better interpretability compared with modern CNN methods. * The work was performed while Wenxuan Bao was a visiting student at University of Rochester.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Deep Hierarchical Semantic SegmentationLiulei Li, Tianfei Zhou, Wenguan Wang, Jianwu Li 等CVPR 2022 · 被引用 181 次
- LogicSeg: Parsing Visual Semantics with Neural Logic Learning and ReasoningLiulei Li, Wenguan Wang, Yang YiICCV 2023 · 被引用 52 次
- Flattening the Parent Bias: Hierarchical Semantic Segmentation in the Poincaré BallSimon Weber, Baris Zöngür, Nikita Araslanov, Daniel CremersCVPR 2024
- Visual Dependency Transformers: Dependency Tree Emerges from Reversed AttentionMingyu Ding, Yikang Shen, Lijie Fan, Zhenfang Chen 等CVPR 2023
它引用的顶会 Paper3
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 被引用 555 次
- Adaptive Context Network for Scene ParsingJun Fu, Jing Liu, Yuhang Wang, Yong Li 等ICCV 2019 · 被引用 148 次
相关 Paper
- Exploring Figure-Ground Assignment Mechanism in Perceptual OrganizationWei Zhai, Yang Cao, Jing Zhang, Zheng-Jun ZhaNeurIPS 2022 · 被引用 36 次
- Hierarchical Human Parsing With Typed Part-Relation ReasoningWenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang 等CVPR 2020
- Learning Compositional Neural Information Fusion for Human ParsingWenguan Wang, Zhijie Zhang, Siyuan Qi, Jianbing Shen 等ICCV 2019 · 被引用 131 次
- Perceptual Group Tokenizer: Building Perception with Iterative GroupingZhiwei Deng, Ting Chen, Yang LiICLR 2024 · 被引用 4 次
- Multi-modal Grouping Network for Weakly-Supervised Audio-Visual Video ParsingShentong Mo, Yapeng TianNeurIPS 2022 · 被引用 73 次
