TACO: Taming Diffusion for In-the-Wild Video Amodal Completion
Ruijie Lu, Yixin Chen, Yu Liu, Jiaxiang Tang, Junfeng Ni, Diwen Wan, Gang Zeng, Siyuan Huang
Abstract
Humans can infer complete shapes and appearances of objects from limited visual cues, relying on extensive prior knowledge of the physical world. However, completing partially observable objects while ensuring consistency across video frames remains challenging for existing models, especially for unstructured, in-the-wild videos. This paper tackles Video Amodal Completion (VAC), aiming to generate the complete object consistently throughout the video given a visual prompt specifying the object of interest. Leveraging the rich, consistent manifolds learned by pre-trained video diffusion models, we propose a conditional diffusion model, TACO, that repurposes these manifolds for VAC. To enable its effective and robust generalization to challenging in-thewild scenarios, we curate a large-scale synthetic dataset with multiple difficulty levels by systematically imposing occlusions onto un-occluded videos. Building on this, we devise a progressive fine-tuning paradigm that starts with simpler recovery tasks and gradually advances to more complex ones. We demonstrate TACO's versatility on a wide range of in-the-wild videos from Internet, as well as on diverse, unseen datasets commonly used in autonomous driving, robotic manipulation, and scene understanding. Moreover, we show that TACO can be effectively applied to various downstream tasks like object reconstruction and pose estimation, highlighting its potential to facilitate physical world understanding and reasoning. Our project page is available at https://jason-aplp.github.io/TACO/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28aa5c18-cb56-4d54-a113-873ebaf4c88aCited by top-tier papers6
- Trace3D: Consistent Segmentation Lifting via Gaussian Instance TracingHongyu Shen, Junfeng Ni, Yixin Chen, Weishuo Li et al.ICCV 2025 · 5 citations
- SynergyAmodal: Deocclude Anything with Text ControlXinyang Li, Chengjie Yi, Jiawei Lai, Mingbao Lin et al.ACM MM 2025 · 3 citations
- GWM: Towards Scalable Gaussian World Models for Robotic ManipulationGuanxing Lu, Baoxiong Jia, Puhao Li, Yixin Chen et al.ICCV 2025 · 1 citation
- Lifting Unlabeled Internet-level Data for 3D Scene UnderstandingYixin Chen, Yaowei Zhang, Huangyue Yu, Junchao He et al.CVPR 2026 · 1 citation
- METASCENES: Towards Automated Replica Creation for Real-world 3D ScansHuangyue Yu, Baoxiong Jia, Yixin Chen, Yandan Yang et al.CVPR 2025
Builds on60
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Image Guides Images: Consistent Video Amodal Completion with Rectified In-Context Exemplar GuidanceXiaoyu Kong, Ketong Ren, Dongyu She, Weiming Dong et al.CVPR 2026
- Object-level Scene DeocclusionZhengzhe Liu, Qing Liu, Chirui Chang, Jianming Zhang et al.SIGGRAPH 2024 · 9 citations
- Using Diffusion Priors for Video Amodal SegmentationKaihua Chen, Deva Ramanan, Tarasha KhuranaCVPR 2025
- Amodal Completion via Progressive Mixed Context DiffusionKatherine Xu, Lingzhi Zhang, Jianbo ShiCVPR 2024 · 20 citations
- Amodal Ground Truth and Completion in the WildGuanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew ZissermanCVPR 2024 · 23 citations
