Variational Context-Deformable ConvNets for Indoor Scene Parsing
Zhitong Xiong, Yuan Yuan, Nianhui Guo, Qi Wang
Abstract
Context information is critical for image semantic segmentation. Especially in indoor scenes, the large variation of object scales makes spatial-context an important factor for improving the segmentation performance. Thus, in this paper, we propose a novel variational context-deformable (VCD) module to learn adaptive receptive-field in a structured fashion. Different from standard ConvNets, which share fixed-size spatial context for all pixels, the VCD module learns a deformable spatial-context with the guidance of depth information: depth information provides clues for identifying real local neighborhoods. Specifically, adaptive Gaussian kernels are learned with the guidance of multimodal information. By multiplying the learned Gaussian kernel with standard convolution filters, the VCD module can aggregate flexible spatial context for each pixel during convolution. The main contributions of this work are as follows: 1) a novel VCD module is proposed, which exploits learnable Gaussian kernels to enable feature learning with structured adaptive-context; 2) variational Bayesian probabilistic modeling is introduced for the training of VCD module, which can make it continuous and more stable; 3) a perspective-aware guidance module is designed to take advantage of multi-modal information for RGB-D segmentation. We evaluate the proposed approach on three widelyused datasets, and the performance improvement has shown the effectiveness of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 368cd558-1e7d-4736-afc4-5b352fc460beCited by top-tier papers5
- Omni-Scale CNNs: a simple and effective kernel size configuration for time series classificationWensi Tang, Guodong Long, Lu Liu, Tianyi Zhou et al.ICLR 2022 · 163 citations
- FlexConv: Continuous Kernel Convolutions With Differentiable Kernel SizesDavid W. Romero, Robert-Jan Bruintjes, Jakub Mikolaj Tomczak, Erik J. Bekkers et al.ICLR 2022 · 94 citations
- Dynamic Sparse Network for Time Series Classification: Learning What to "See"Qiao Xiao, Boqian Wu, Yu Zhang, Shiwei Liu et al.NeurIPS 2022 · 45 citations
- Deep Continuous NetworksNergis Tomen, Silvia-Laura Pintea, Jan van GemertICML 2021 · 15 citations
- EffConv: Efficient Learning of Kernel Sizes for Convolution Layers of CNNsAlireza Ganjdanesh, Shangqian Gao, Heng HuangAAAI 2023 · 11 citations
Builds on2
Related papers
- Dynamic Sampling Network for Semantic SegmentationBin Fu, Junjun He, Zhengfu Zhang, Yu QiaoAAAI 2020 · 6 citations
- Anisotropic Convolutional Networks for 3D Semantic Scene CompletionJie Li, Kai Han, Peng Wang, Yu Liu et al.CVPR 2020
- Dynamic Multi-Scale Filters for Semantic SegmentationJunjun He, Zhongying Deng, Yu QiaoICCV 2019 · 287 citations
- SCF-Net: Learning Spatial Contextual Features for Large-Scale Point Cloud SegmentationSiqi Fan, Qiulei Dong, Fenghua Zhu, Yisheng Lv et al.CVPR 2021
- MLCVNet: Multi-Level Context VoteNet for 3D Object DetectionQian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang et al.CVPR 2020
