STREAMER: Streaming Representation Learning and Event Segmentation in a Hierarchical Manner
Ramy Mounir, Sujal Vijayaraghavan, Sudeep Sarkar
Abstract
We present a novel self-supervised approach for hierarchical representation learning and segmentation of perceptual inputs in a streaming fashion. Our research addresses how to semantically group streaming inputs into chunks at various levels of a hierarchy while simultaneously learning, for each chunk, robust global representations throughout the domain. To achieve this, we propose STREAMER, an architecture that is trained layer-by-layer, adapting to the complexity of the input domain. In our approach, each layer is trained with two primary objectives: making accurate predictions into the future and providing necessary information to other levels for achieving the same objective. The event hierarchy is constructed by detecting prediction error peaks at different levels, where a detected boundary triggers a bottom-up information flow. At an event boundary, the encoded representation of inputs at one layer becomes the input to a higher-level layer. Additionally, we design a communication module that facilitates top-down and bottom-up exchange of information during the prediction process. Notably, our model is fully self-supervised and trained in a streaming manner, enabling a single pass on the training data. This means that the model encounters each input only once and does not store the data. We evaluate the performance of our model on the egocentric EPIC-KITCHENS dataset, specifically focusing on temporal event segmentation. Furthermore, we conduct event retrieval experiments using the learned representations to demonstrate the high quality of our video event representations. Illustration videos and code are available on our project page: https://ramymounir.com/publications/streamer . Cross-layer communication CNN encoder CNN encoder CNN decoder
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 275d9c38-e5a7-4ea3-8105-66d847d4e0bcCited by top-tier papers4
- Hierarchical Vector Quantization for Unsupervised Action SegmentationFederico Spurio, Emad Bahrami, Gianpiero Francesca, Juergen GallAAAI 2025 · 17 citations
- Representing Part-Whole Hierarchies in Foundation Models by Learning Localizability, Composability, and Decomposability from Anatomy via Self-SupervisionMohammad Reza Hosseinzadeh Taher, Michael B. Gotway, Jianming LiangCVPR 2024 · 12 citations
- Predictive Attractor ModelsRamy Mounir, Sudeep SarkarNeurIPS 2024 · 2 citations
- Hierarchical Action Learning for Weakly-Supervised Action SegmentationJunxian Huang, Ruichu Cai, Juntao Fang, Hao Zhu et al.CVPR 2026 · 1 citation
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- Learning State-Aware Visual Representations from Audible InteractionsHimangi Mittal, Pedro Morgado, Unnat Jain, Abhinav GuptaNeurIPS 2022 · 30 citations
- Open-Ended Hierarchical Streaming Video Understanding with Vision Language ModelsHyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi et al.ICCV 2025 · 2 citations
- Generative Hybrid Representations for Activity Forecasting With No-Regret LearningJiaqi Guan, Ye Yuan, Kris M. Kitani, Nicholas RhinehartCVPR 2020
- Moving Off-the-Grid: Scene-Grounded Video RepresentationsSjoerd van Steenkiste, Daniel Zoran, Yi Yang, Yulia Rubanova et al.NeurIPS 2024 · 13 citations
- Variational Predictive Routing with Nested Subjective TimescalesAlexey Zakharov, Qinghai Guo, Zafeirios FountasICLR 2022 · 12 citations
