CondDiff-AMO: Integrating Conditional Diffusion Mechanism for Unified Amodal Mask Generation
Caijie Zhao, Bob Zhang
Abstract
Aiming to estimate the full extent of partially occluded objects, amodal segmentation is a critical capability for visual intelligence. Existing methods suffer from limitations in efficiency and precision, due to their reliance on auxiliary information or two-stage architectures. Furthermore, they lack generalizability, failing to meet practical requirements. To overcome these challenges, we proposed a new paradigm, CondDiff-AMO, that interprets amodal segmentation as a denoising problem by leveraging diffusion models. Methodologically, the designed novel framework consists of three key innovations to adapt the task characteristics and unlocks the diffusion models’ potential in amodal segmentation, including a masking strategy in the forward process, an adaptive transformer for conditional feature extraction, and visual-guided sampling. In the forward process, progressive masking strategy converts ground-truth masks to visible masks, simulating amodal segmentation process to enhance reasoning regarding occluded areas. For architectural design, a pyramid network with feature refinement extracts adaptive and representative conditional priors, improving the guidance in the denoising process of diffusion models. As for the sampling stage, a visible mask is incorporated with an ensemble strategy, restricting the prediction on occluded part. Experiments were conducted on five well-known datasets under supervised and zero-shot learning, with the results confirming that CondDiff-AMO outperforms state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Label-Efficient Semantic Segmentation with Diffusion ModelsDmitry Baranchuk, Andrey Voynov, Ivan Rubachev, Valentin Khrulkov et al.ICLR 2022 · 700 citations
- MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with TransformerJunde Wu, Wei Ji, Huazhu Fu, Min Xu et al.AAAI 2024 · 311 citations
- Amodal Segmentation Based on Visible Region Segmentation and Shape PriorYuting Xiao, Yanyu Xu, Ziming Zhong, Weixin Luo et al.AAAI 2021 · 76 citations
Related papers
- pix2gestalt: Amodal Segmentation by Synthesizing WholesEge Ozguroglu, Ruoshi Liu, Dídac Surís, Dian Chen et al.CVPR 2024 · 24 citations
- Tuning-Free Amodal Segmentation via the Occlusion-Free Bias of Inpainting ModelsJae Joong Lee, Bedrich Benes, Raymond A. YehAAAI 2026 · 2 citations
- Amodal Completion via Progressive Mixed Context DiffusionKatherine Xu, Lingzhi Zhang, Jianbo ShiCVPR 2024 · 20 citations
- CamoDiffusion: Camouflaged Object Detection via Conditional Diffusion ModelsZhongxi Chen, Ke Sun, Xianming LinAAAI 2024 · 61 citations
- Towards Efficient Foundation Model for Zero-shot Amodal SegmentationZhaochen Liu, Limeng Qiao, Xiangxiang Chu, Lin Ma et al.CVPR 2025
