pix2gestalt: Amodal Segmentation by Synthesizing Wholes
Ege Ozguroglu, Ruoshi Liu, Dídac Surís, Dian Chen, Achal Dave, Pavel Tokmakov, Carl Vondrick
Abstract
We introduce pix2gestalt, a framework for zero-shot amodal segmentation, which learns to estimate the shape and appearance of whole objects that are only partially visible behind occlusions. By capitalizing on large-scale diffusion models and transferring their representations to this task, we learn a conditional diffusion model for reconstructing whole objects in challenging zero-shot cases, including examples that break natural and physical priors, such as art. As training data, we use a synthetically curated dataset containing occluded objects paired with their whole counterparts. Experiments show that our approach outperforms supervised baselines on established benchmarks. Our model can furthermore be used to significantly improve the performance of existing object recognition and 3D reconstruction methods in the presence of occlusions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59a1a87f-2e33-4fb5-84f9-4d903d0dea0cCited by top-tier papers36
- HoloPart: Generative 3D Part Amodal SegmentationYunhan Yang, Yuanchen Guo, Yukun Huang, Zi-Xin Zou et al.ICLR 2026 · 62 citations
- Amodal Ground Truth and Completion in the WildGuanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew ZissermanCVPR 2024 · 23 citations
- Amodal Completion via Progressive Mixed Context DiffusionKatherine Xu, Lingzhi Zhang, Jianbo ShiCVPR 2024 · 20 citations
- Visual Jenga: Discovering Object Dependencies via Counterfactual InpaintingAnand Bhattad, Konpat Preechakul, Alexei A. EfrosNeurIPS 2025 · 13 citations
- ReferevErything: Towards Segmenting Everything we can Speak of in VideosAnurag Bagchi, Zhipeng Bao, Yu-Xiong Wang, Pavel Tokmakov et al.ICCV 2025 · 11 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Tuning-Free Amodal Segmentation via the Occlusion-Free Bias of Inpainting ModelsJae Joong Lee, Bedrich Benes, Raymond A. YehAAAI 2026 · 2 citations
- CondDiff-AMO: Integrating Conditional Diffusion Mechanism for Unified Amodal Mask GenerationCaijie Zhao, Bob ZhangAAAI 2026
- Conditional Latent Diffusion Models for Zero-Shot Instance SegmentationMaximilian Ulmer, Wout Boerdijk, Rudolph Triebel, Maximilian DurnerICCV 2025 · 1 citation
- Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask CompletionBowen Zhang, Qing Liu, Jianming Zhang, Yilin Wang et al.AAAI 2024 · 4 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
