Unique Lives, Shared World: Learning from Single-Life Videos
Tengda Han, Sayna Ebrahimi, Dilara Gokay, Li Yang Ku, Maks Ovsjanikov, Iva Babukova, Daniel Zoran, Viorica Patraucean, João Carreira, Andrew Zisserman, Dima Damen
Abstract
We introduce the "single-life" learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual. We leverage the multiple viewpoints naturally captured within a single life to learn a visual encoder in a self-supervised manner. Our experiments demonstrate three key findings. First, models trained independently on different lives develop a highly aligned geometric understanding. We demonstrate this by training visual encoders on distinct datasets each capturing a different life, both indoors and outdoors, as well as introducing a novel cross-attention-based metric to quantify the functional alignment of the internal representations developed by different models. Second, we show that single-life models learn generalizable geometric representations that effectively transfer to downstream tasks, such as depth estimation, in unseen environments. Third, we demonstrate that training on up to 30 hours from one week of the same person's life leads to comparable performance to training on 30 hours of diverse web data, highlighting the strength of single-life representation learning. Overall, our results establish that the shared structure of the world, both leads to consistency in models trained on individual lives, and provides a powerful signal for visual representation learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1409654b-8c41-41dc-98c1-6f2dccb94caaBuilds on22
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 356 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- Revisiting Model Stitching to Compare Neural RepresentationsYamini Bansal, Preetum Nakkiran, Boaz BarakNeurIPS 2021 · 253 citations
Related papers
- SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionBehzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe ThiranICCV 2019 · 44 citations
- LightDepth: Single-View Depth Self-Supervision from Illumination DeclineJavier Rodriguez Puigvert, Victor M. Batlle, J. M. M. Montiel, Ruben Martinez-Cantin et al.ICCV 2023 · 22 citations
- Learning by Aligning Videos in TimeSanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed et al.CVPR 2021
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingQitao Zhao, Hao Tan, Qianqian Wang, Sai Bi et al.CVPR 2026 · 24 citations
- Multiview Compressive Coding for 3D ReconstructionChao-Yuan Wu, Justin Johnson, Jitendra Malik, Christoph Feichtenhofer et al.CVPR 2023
