Tracking Through Containers and Occluders in the Wild
Basile Van Hoorick, Pavel Tokmakov, Simon Stent, Jie Li, Carl Vondrick
Abstract
Tracking objects with persistence in cluttered and dynamic environments remains a difficult challenge for computer vision systems. In this paper, we introduce TCOW, a new benchmark and model for visual tracking through heavy occlusion and containment. We set up a task where the goal is to, given a video sequence, segment both the projected extent of the target object, as well as the surrounding container or occluder whenever one exists. To study this task, we create a mixture of synthetic and annotated real datasets to support both supervised learning and structured evaluation of model performance under various forms of task variation, such as moving or nested containment. We evaluate two recent transformer-based video models and find that while they can be surprisingly capable of tracking targets under certain settings of task variation, there remains a considerable performance gap before we can claim a tracking model to have acquired a true notion of object permanence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4efd084d-82be-4f01-a10b-e1350a5556d1Cited by top-tier papers6
- Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under OcclusionsRuihai Wu, Kai Cheng, Yan Zhao, Chuanruo Ning et al.NeurIPS 2023 · 43 citations
- Amodal Ground Truth and Completion in the WildGuanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew ZissermanCVPR 2024 · 23 citations
- TACO: Taming Diffusion for In-the-Wild Video Amodal CompletionRuijie Lu, Yixin Chen, Yu Liu, Jiaxiang Tang et al.ICCV 2025 · 3 citations
- Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view DiffusionZeren Xiong, Zikun Chen, Zedong Zhang, Xiang Li et al.ACM MM 2025 · 2 citations
- Temporally Consistent Object-Centric Learning by Contrasting SlotsAnna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius et al.CVPR 2025
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
Related papers
- Transparent Object Tracking BenchmarkHeng Fan, Halady Akhilesha Miththanthaya, Harshit, Siranjiv Ramana Rajan et al.ICCV 2021 · 32 citations
- Hopper: Multi-hop Transformer for Spatiotemporal ReasoningHonglu Zhou, Asim Kadav, Farley Lai, Alexandru Niculescu-Mizil et al.ICLR 2021 · 19 citations
- Opening up Open World TrackingYang Liu, Idil Esen Zulfikar, Jonathon Luiten, Achal Dave et al.CVPR 2022 · 40 citations
- Segmenting Moving Objects via an Object-Centric Layered RepresentationJunyu Xie, Weidi Xie, Andrew ZissermanNeurIPS 2022 · 74 citations
- Video OWL-ViT: Temporally-consistent open-world localization in videoGeorg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic et al.ICCV 2023 · 22 citations
