Deep Grouping Model for Unified Perceptual Parsing
Zhiheng Li, Wenxuan Bao, Jiayang Zheng, Chenliang Xu
Abstract
The perceptual-based grouping process produces a hierarchical and compositional image representation that helps both human and machine vision systems recognize heterogeneous visual concepts. Examples can be found in the classical hierarchical superpixel segmentation or image parsing works. However, the grouping process is largely overlooked in modern CNN-based image segmentation networks due to many challenges, including the inherent incompatibility between the grid-shaped CNN feature map and the irregularshaped perceptual grouping hierarchy. Overcoming these challenges, we propose a deep grouping model (DGM) that tightly marries the two types of representations and defines a bottom-up and a top-down process for feature exchanging. When evaluating the model on the recent Broden+ dataset for the unified perceptual parsing task, it achieves state-ofthe-art results while having a small computational overhead compared to other contextual-based segmentation models. Furthermore, the DGM has better interpretability compared with modern CNN methods. * The work was performed while Wenxuan Bao was a visiting student at University of Rochester.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Deep Hierarchical Semantic SegmentationLiulei Li, Tianfei Zhou, Wenguan Wang, Jianwu Li et al.CVPR 2022 · 181 citations
- LogicSeg: Parsing Visual Semantics with Neural Logic Learning and ReasoningLiulei Li, Wenguan Wang, Yang YiICCV 2023 · 52 citations
- Flattening the Parent Bias: Hierarchical Semantic Segmentation in the Poincaré BallSimon Weber, Baris Zöngür, Nikita Araslanov, Daniel CremersCVPR 2024
- Visual Dependency Transformers: Dependency Tree Emerges from Reversed AttentionMingyu Ding, Yikang Shen, Lijie Fan, Zhenfang Chen et al.CVPR 2023
Builds on3
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 555 citations
- Adaptive Context Network for Scene ParsingJun Fu, Jing Liu, Yuhang Wang, Yong Li et al.ICCV 2019 · 148 citations
Related papers
- Exploring Figure-Ground Assignment Mechanism in Perceptual OrganizationWei Zhai, Yang Cao, Jing Zhang, Zheng-Jun ZhaNeurIPS 2022 · 36 citations
- Hierarchical Human Parsing With Typed Part-Relation ReasoningWenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang et al.CVPR 2020
- Learning Compositional Neural Information Fusion for Human ParsingWenguan Wang, Zhijie Zhang, Siyuan Qi, Jianbing Shen et al.ICCV 2019 · 131 citations
- Perceptual Group Tokenizer: Building Perception with Iterative GroupingZhiwei Deng, Ting Chen, Yang LiICLR 2024 · 4 citations
- Multi-modal Grouping Network for Weakly-Supervised Audio-Visual Video ParsingShentong Mo, Yapeng TianNeurIPS 2022 · 73 citations
