Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners
Nikita Araslanov, Martin Sundermeyer, Hidenobu Matsuki, David Joseph Tan, Federico Tombari
Abstract
One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of visual scenes at the pixel level. Existing frameworks either train on image-based pretext tasks, which do not account for dynamic elements, or on video sequences for action-level reasoning, which does not scale to dense pixel-level prediction. We present a framework that learns pixel-accurate feature descriptors from videos, LILA. The core element of our training framework is linear in-context learning. LILA leverages spatio-temporal cue maps -- depth and motion -- estimated with off-the-shelf networks. Despite the noisy nature of those cues, LILA trains effectively on uncurated video datasets, embedding semantic and geometric properties in a temporally consistent manner. We demonstrate compelling empirical benefits of the learned representation across a diverse suite of vision tasks: video object segmentation, surface normal estimation and semantic segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 200bbc65-6d3c-4d33-9e07-db17c3878a7aBuilds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Look, Remember and Reason: Grounded Reasoning in Videos with Language ModelsApratim Bhattacharyya, Sunny Panchal, Reza Pourreza, Mingu Lee et al.ICLR 2024 · 15 citations
- EVAL: Explainable Video Anomaly LocalizationAshish Singh, Michael J. Jones, Erik G. Learned-MillerCVPR 2023
- Does Visual Pretraining Help End-to-End Reasoning?Chen Sun, Calvin Luo, Xingyi Zhou, Anurag Arnab et al.NeurIPS 2023 · 4 citations
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 59 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
