Learning from One Continuous Video Stream
João Carreira, Michael King, Viorica Patraucean, Dilara Gokay, Catalin Ionescu, Yi Yang, Daniel Zoran, Joseph Heyward, Carl Doersch, Yusuf Aytar, Dima Damen, Andrew Zisserman
摘要
We introduce a framework for online learning from a single continuous video stream -the way people and animals learn, without mini-batches, data augmentation or shuffling. This poses great challenges given the high correlation between consecutive video frames and there is very little prior work on it. Our framework allows us to do a first deep dive into the topic and includes a collection of streams and tasks composed from two existing video datasets, plus methodology for performance evaluation that considers both adaptation and generalization. We employ pixel-to-pixel modelling as a practical and flexible way to switch between pre-training and single-stream evaluation as well as between arbitrary tasks, without ever requiring changes to models and always using the same pixel loss. Equipped with this framework we obtained large singlestream learning gains from pre-training with a novel family of future prediction tasks, found that momentum hurts, and that the pace of weight updates matters. The combination of these insights leads to matching the performance of IID learning with batch size 1, when using the same architecture and without costly replay buffers. An overview of the paper is available online at https://sites.google . com/view/one-stream-video.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Recurrent Video Masked AutoencodersDaniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson 等CVPR 2026 · 被引用 9 次
- Unique Lives, Shared World: Learning from Single-Life VideosTengda Han, Sayna Ebrahimi, Dilara Gokay, Li Yang Ku 等CVPR 2026 · 被引用 2 次
- Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video CaptioningYolo Yunlong Tang, Chao Huang, Susan Liang, Jing Bi 等CVPR 2026 · 被引用 2 次
- Mosic: Optimal-Transport Motion Trajectory for Dense Self-Supervised LearningMohammadreza Salehi, Shashanka Venkataramanan, Ioana Simion, Efstratios Gavves 等ICCV 2025
- Learning from Streaming Video with Orthogonal GradientsTengda Han, Dilara Gokay, Joseph Heyward, Chuhan Zhang 等CVPR 2025
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- PIVOT: Prompting for Video Continual LearningAndrés Villa, Juan León Alcázar, Motasem Alfarra, Kumail Alhamoud 等CVPR 2023
- Label-Efficient Online Continual Object Detection in Streaming VideoJay Zhangjie Wu, David Junhao Zhang, Wynne Hsu, Mengmi Zhang 等ICCV 2023 · 被引用 24 次
- Sideways: Depth-Parallel Training of Video ModelsMateusz Malinowski, Grzegorz Swirszcz, João Carreira, Viorica PatrauceanCVPR 2020
- How Well Does Self-Supervised Pre-Training Perform with Streaming Data?Dapeng Hu, Shipeng Yan, Qizhengqiu Lu, Lanqing Hong 等ICLR 2022 · 被引用 36 次
- FlashDepth: Real-Time Streaming Video Depth Estimation at 2K ResolutionGene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah 等ICCV 2025 · 被引用 1 次
