Learning from Streaming Video with Orthogonal Gradients
Tengda Han, Dilara Gokay, Joseph Heyward, Chuhan Zhang, Daniel Zoran, Viorica Patraucean, João Carreira, Dima Damen, Andrew Zisserman
Abstract
We address the challenge of representation learning from a continuous stream of video as input, in a self-supervised manner. This differs from the standard approaches to video learning where videos are chopped and shuffled during training in order to create a non-redundant batch that satisfies the independently and identically distributed (IID) sample assumption expected by conventional training paradigms. When videos are only available as a continuous stream of input, the IID assumption is evidently broken, leading to poor performance. We demonstrate the drop in performance when moving from shuffled to sequential learning on three tasks: the one-video representation learning method DoRA, standard VideoMAE on multi-video datasets, and the task of future video prediction. To address this drop, we propose a geometric modification to standard optimizers, to decorrelate batches by utilising orthogonal gradients during training. The proposed modification can be applied to any optimizer -we demonstrate it with Stochastic Gradient Descent (SGD) and AdamW. Our proposed orthogonal optimizer allows models trained from streaming videos to alleviate the drop in representation learning performance, as evaluated on downstream tasks. On three scenarios (DoRA, VideoMAE, future prediction), we show our orthogonal optimizer outperforms the strong AdamW in all three scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Unique Lives, Shared World: Learning from Single-Life VideosTengda Han, Sayna Ebrahimi, Dilara Gokay, Li Yang Ku et al.CVPR 2026 · 2 citations
- Learning Streaming Video Representation via Multitask TrainingYibin Yan, Jilan Xu, Shangzhe Di, Yikun Liu et al.ICCV 2025 · 1 citation
- Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement LearningNilaksh, Antoine Clavaud, Mathieu Reymond, Francois Rivest et al.ICML 2026
Builds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
Related papers
- Moving Off-the-Grid: Scene-Grounded Video RepresentationsSjoerd van Steenkiste, Daniel Zoran, Yi Yang, Yulia Rubanova et al.NeurIPS 2024 · 13 citations
- MGMAE: Motion Guided Masking for Video Masked AutoencodingBingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao et al.ICCV 2023 · 58 citations
- Learning from One Continuous Video StreamJoão Carreira, Michael King, Viorica Patraucean, Dilara Gokay et al.CVPR 2024
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- Learning by Aligning Videos in TimeSanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed et al.CVPR 2021
