IDRNet: Intervention-Driven Relation Network for Semantic Segmentation
Zhenchao Jin, Xiaowei Hu, Lingting Zhu, Luchuan Song, Li Yuan, Lequan Yu
摘要
Co-occurrent visual patterns suggest that pixel relation modeling facilitates dense prediction tasks, which inspires the development of numerous context modeling paradigms, e.g., multi-scale-driven and similarity-driven context schemes. Despite the impressive results, these existing paradigms often suffer from inadequate or ineffective contextual information aggregation due to reliance on large amounts of predetermined priors. To alleviate the issues, we propose a novel Intervention-Driven Relation Network (IDRNet), which leverages a deletion diagnostics procedure to guide the modeling of contextual relations among different pixels. Specifically, we first group pixel-level representations into semantic-level representations with the guidance of pseudo labels and further improve the distinguishability of the grouped representations with a feature enhancement module. Next, a deletion diagnostics procedure is conducted to model relations of these semantic-level representations via perceiving the network outputs and the extracted relations are utilized to guide the semantic-level representations to interact with each other. Finally, the interacted representations are utilized to augment original pixel-level representations for final predictions. Extensive experiments are conducted to validate the effectiveness of IDRNet quantitatively and qualitatively. Notably, our intervention-driven context scheme brings consistent performance improvements to state-of-the-art segmentation frameworks and achieves competitive results on popular benchmark datasets, including ADE20K, COCO-Stuff, PASCAL-Context, LIP, and Cityscapes. Code is available at https://github.com/SegmentationBLWX/sssegmentation . Table 1: Performance comparison between existing context schemes and our IDRNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any ScenarioYuhan Li, Hao Zhou, Wenxiang Shang, Ran Lin 等NeurIPS 2024 · 被引用 31 次
- AUCSeg: AUC-oriented Pixel-level Long-tail Semantic SegmentationBoyu Han, Qianqian Xu, Zhiyong Yang, Shilong Bao 等NeurIPS 2024 · 被引用 26 次
- AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion ModelsXinghui Li, Qichao Sun, Pengze Zhang, Fulong Ye 等CVPR 2025
它引用的顶会 Paper25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
相关 Paper
- ISNet: Integrate Image-Level and Semantic-Level Context for Semantic SegmentationZhenchao Jin, Bin Liu, Qi Chu, Nenghai YuICCV 2021 · 被引用 88 次
- Context Prior for Scene SegmentationChangqian Yu, Jingbo Wang, Changxin Gao, Gang Yu 等CVPR 2020
- RANet: Region Attention Network for Semantic SegmentationDingguo Shen, Yuanfeng Ji, Ping Li, Yi Wang 等NeurIPS 2020 · 被引用 43 次
- SGINet: Toward Sufficient Interaction Between Single Image Deraining and Semantic SegmentationYanyan Wei, Zhao Zhang, Huan Zheng, Richang Hong 等ACM MM 2022 · 被引用 34 次
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang 等CVPR 2022 · 被引用 228 次
