A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives
Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, Giuseppe Averta
Abstract
Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once. We believe that - to effectively transfer such an holistic perception to intelligent machines - an important role is played by learning to correlate concepts and to abstract knowledge coming from different tasks, to synergistically exploit them when learning novel skills. To accomplish this, we look for a unified approach to video understanding which combines shared temporal modelling of human actions with minimal overhead, to support multiple downstream tasks and enable cooperation when learning novel skills. We then propose EgoPack, a solution that creates a collection of task perspectives that can be carried across downstream tasks and used as a potential source of additional insights, as a backpack of skills that a robot can carry around and use when needed. We demonstrate the effectiveness and efficiency of our approach on four Ego4D benchmarks, outperforming current state-of-the-art methods. Project webpage: sapeirone.github.io/EgoPack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsYuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu et al.NeurIPS 2024 · 31 citations
- Video2BEV: Transforming Drone Videos to BEVs for Video-Based Geo-LocalizationHao Ju, Shaofei Huang, Si Liu, Zhedong ZhengICCV 2025 · 5 citations
- EgoLife: Towards Egocentric Life AssistantJingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong et al.CVPR 2025
- HiERO: Understanding the Hierarchy of Human Behavior Enhances Reasoning on Egocentric VideosSimone Alberto Peirone, Francesca Pistilli, Giuseppe AvertaICCV 2025
Builds on31
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying et al.ICML 2020 · 1,439 citations
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas et al.ICML 2020 · 651 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- Egocentric Video Task TranslationZihui Xue, Yale Song, Kristen Grauman, Lorenzo TorresaniCVPR 2023
- CEL: Continual Ego, Exo, and Ego-Exo LearningHongwei Yan, Kanglei Zhou, Yuchen Liu, Qingyu Shi et al.ICML 2026
- Shaping embodied agent behavior with activity-context priors from egocentric videoTushar Nagarajan, Kristen GraumanNeurIPS 2021 · 23 citations
- EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real WorldYifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang et al.CVPR 2024
- EgoEnv: Human-centric environment representations from egocentric videoTushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta Desai, James Hillis et al.NeurIPS 2023 · 28 citations
