Multi-View Aggregation Network for Dichotomous Image Segmentation
Qian Yu, Xiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu
Abstract
Dichotomous Image Segmentation (DIS) has recently emerged towards high-precision object segmentation from high-resolution natural images. When designing an effective DIS model, the main challenge is how to balance the semantic dispersion of high-resolution targets in the small receptive field and the loss of high-precision details in the large receptive field. Existing methods rely on tedious multiple encoder-decoder streams and stages to gradually complete the global localization and local refinement. Human visual system captures regions of interest by observing them from multiple views. Inspired by it, we model DIS as a multi-view object perception problem and provide a parsi-monious multi-view aggregation network (MVANet), which unifies the feature fusion of the distant view and close-up view into a single stream with one encoder-decoder structure. With the help of the proposed multi-view complementary localization and refinement modules, our approach established long-range, profound visual interactions across multiple views, allowing the features of the detailed close-up view to focus on highly slender structures. Experiments on the popular DIS-5K dataset show that our MVANet significantly outperforms state-of-the-art methods in both accuracy and speed. The source code and datasets will be publicly available at MVANet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d79df926-8e50-4f94-9d69-784cf25983d3Cited by top-tier papers13
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng et al.NeurIPS 2025 · 45 citations
- S3OD: Towards Generalizable Salient Object Detection with Synthetic DataOrest Kupyn, Hirokatsu Kataoka, Christian RupprechtICLR 2026 · 6 citations
- Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene ModelingQirui Wu, Denys Iliash, Daniel Ritchie, Manolis Savva et al.ICCV 2025 · 4 citations
- Towards Fine-Grained Interactive Segmentation in Images and VideosYuan Yao, Qiushi Yang, Miaomiao Cui, Liefeng BoICCV 2025 · 2 citations
- Fix False Transparency by Noise Guided SplattingAly El Hakie, Yiren Lu, Yu Yin, Michael Jenkins et al.NeurIPS 2025 · 2 citations
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang et al.ICCV 2019 · 694 citations
- Global Context-Aware Progressive Aggregation Network for Salient Object DetectionZuyao Chen, Qianqian Xu, Runmin Cong, Qingming HuangAAAI 2020 · 481 citations
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
Related papers
- Unite-Divide-Unite: Joint Boosting Trunk and Structure for High-accuracy Dichotomous Image SegmentationJialun Pei, Zhangjun Zhou, Yueming Jin, He Tang et al.ACM MM 2023 · 21 citations
- LawDIS: Language-Window-Based Controllable Dichotomous Image SegmentationXinyu Yan, Meijun Sun, Ge-Peng Ji, Fahad Shahbaz Khan et al.ICCV 2025 · 3 citations
- High-Precision Dichotomous Image Segmentation via Depth Integrity-Prior and Fine-Grained Patch StrategyXianjie Liu, Keren Fu, Qijun ZhaoCVPR 2026 · 2 citations
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
- DualDis: A Dual Disentanglement Network for Vehicle Re-identificationWenying He, Feiyu Wang, Guangquan Xu, Yude Bai et al.WWW 2026
