Image Guides Images: Consistent Video Amodal Completion with Rectified In-Context Exemplar Guidance
Xiaoyu Kong, Ketong Ren, Dongyu She, Weiming Dong, Miao Wang
摘要
Video amodal completion (VAC) aims to mimic the human brain's ability to implicitly perceive the complete appearance of partially occluded objects, thereby facilitating recognition and understanding. Existing VAC methods finetune video generation models on custom datasets, yet these datasets often have unrealistic distributions and small scales due to the challenges of collecting real amodal data and thus limit their performance and generalization.To address this, we utilize pre-trained image inpainting models for VAC and introduce in-context (IC) learning to enhance inter-frame consistency. However, despite the satisfactory performance of DiT-based IC Learning in generation tasks, task-agnostic global information often utilizes irrelevant scene information, resulting in completion failures when applied to amodal completion task. Additionally, IC Learning faces a cold-start problem with the exemplar construction. To this end, we propose a consistency video amodal completion with rectified in-context exemplar guidance. Specifically, we introduce rectified exemplar-guided completion by adjusting the attention weights of exemplar image relative to the target images for consistent completion, and adopt a dual-frame calibrated exemplar rectification to tackle the cold-start issue.Quantitative and qualitative experiments demonstrate that our method outperforms SOTAs, especially in terms of generalization and robustness on uncommon data and under severe occlusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Pix2Video: Video Editing using Image DiffusionDuygu Ceylan, Chun-Hao Paul Huang, Niloy J. MitraICCV 2023 · 被引用 370 次
- SegGPT: Towards Segmenting Everything In ContextXinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang 等ICCV 2023 · 被引用 188 次
相关 Paper
- TACO: Taming Diffusion for In-the-Wild Video Amodal CompletionRuijie Lu, Yixin Chen, Yu Liu, Jiaxiang Tang 等ICCV 2025 · 被引用 3 次
- Amodal Completion via Progressive Mixed Context DiffusionKatherine Xu, Lingzhi Zhang, Jianbo ShiCVPR 2024 · 被引用 20 次
- Using Diffusion Priors for Video Amodal SegmentationKaihua Chen, Deva Ramanan, Tarasha KhuranaCVPR 2025
- BVINet: Unlocking Blind Video Inpainting With Zero AnnotationsZhiliang Wu, Kerui Chen, Kun Li, Hehe Fan 等ICCV 2025 · 被引用 30 次
- Amodal Segmentation through Out-of-Task and Out-of-Distribution Generalization with a Bayesian ModelYihong Sun, Adam Kortylewski, Alan L. YuilleCVPR 2022 · 被引用 26 次
