TACO: Taming Diffusion for In-the-Wild Video Amodal Completion
Ruijie Lu, Yixin Chen, Yu Liu, Jiaxiang Tang, Junfeng Ni, Diwen Wan, Gang Zeng, Siyuan Huang
摘要
Humans can infer complete shapes and appearances of objects from limited visual cues, relying on extensive prior knowledge of the physical world. However, completing partially observable objects while ensuring consistency across video frames remains challenging for existing models, especially for unstructured, in-the-wild videos. This paper tackles Video Amodal Completion (VAC), aiming to generate the complete object consistently throughout the video given a visual prompt specifying the object of interest. Leveraging the rich, consistent manifolds learned by pre-trained video diffusion models, we propose a conditional diffusion model, TACO, that repurposes these manifolds for VAC. To enable its effective and robust generalization to challenging in-thewild scenarios, we curate a large-scale synthetic dataset with multiple difficulty levels by systematically imposing occlusions onto un-occluded videos. Building on this, we devise a progressive fine-tuning paradigm that starts with simpler recovery tasks and gradually advances to more complex ones. We demonstrate TACO's versatility on a wide range of in-the-wild videos from Internet, as well as on diverse, unseen datasets commonly used in autonomous driving, robotic manipulation, and scene understanding. Moreover, we show that TACO can be effectively applied to various downstream tasks like object reconstruction and pose estimation, highlighting its potential to facilitate physical world understanding and reasoning. Our project page is available at https://jason-aplp.github.io/TACO/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Trace3D: Consistent Segmentation Lifting via Gaussian Instance TracingHongyu Shen, Junfeng Ni, Yixin Chen, Weishuo Li 等ICCV 2025 · 被引用 5 次
- SynergyAmodal: Deocclude Anything with Text ControlXinyang Li, Chengjie Yi, Jiawei Lai, Mingbao Lin 等ACM MM 2025 · 被引用 3 次
- GWM: Towards Scalable Gaussian World Models for Robotic ManipulationGuanxing Lu, Baoxiong Jia, Puhao Li, Yixin Chen 等ICCV 2025 · 被引用 1 次
- Lifting Unlabeled Internet-level Data for 3D Scene UnderstandingYixin Chen, Yaowei Zhang, Huangyue Yu, Junchao He 等CVPR 2026 · 被引用 1 次
- METASCENES: Towards Automated Replica Creation for Real-world 3D ScansHuangyue Yu, Baoxiong Jia, Yixin Chen, Yandan Yang 等CVPR 2025
它引用的顶会 Paper60
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Image Guides Images: Consistent Video Amodal Completion with Rectified In-Context Exemplar GuidanceXiaoyu Kong, Ketong Ren, Dongyu She, Weiming Dong 等CVPR 2026
- Object-level Scene DeocclusionZhengzhe Liu, Qing Liu, Chirui Chang, Jianming Zhang 等SIGGRAPH 2024 · 被引用 9 次
- Using Diffusion Priors for Video Amodal SegmentationKaihua Chen, Deva Ramanan, Tarasha KhuranaCVPR 2025
- Amodal Completion via Progressive Mixed Context DiffusionKatherine Xu, Lingzhi Zhang, Jianbo ShiCVPR 2024 · 被引用 20 次
- Amodal Ground Truth and Completion in the WildGuanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew ZissermanCVPR 2024 · 被引用 23 次
