Generative Hybrid Representations for Activity Forecasting With No-Regret Learning
Jiaqi Guan, Ye Yuan, Kris M. Kitani, Nicholas Rhinehart
Abstract
Automatically reasoning about future human behaviors is a difficult problem but has significant practical applications to assistive systems. Part of this difficulty stems from learning systems' inability to represent all kinds of behaviors. Some behaviors, such as motion, are best described with continuous representations, whereas others, such as picking up a cup, are best described with discrete representations. Furthermore, human behavior is generally not fixed: people can change their habits and routines. This suggests these systems must be able to learn and adapt continuously. In this work, we develop an efficient deep generative model to jointly forecast a person's future discrete actions and continuous motions. On a large-scale egocentric dataset, EPIC-KITCHENS, we observe our method generates high-quality and diverse samples while exhibiting better generalization than related generative models. Finally, we propose a variant to continually learn our model from streaming data, observe its practical effectiveness, and theoretically justify its learning efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59e679e7-a0ef-4d03-b03d-26bf45052b30Cited by top-tier papers9
- AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent ForecastingYe Yuan, Xinshuo Weng, Yanglan Ou, Kris KitaniICCV 2021 · 658 citations
- PRECOG: PREdiction Conditioned on Goals in Visual Multi-Agent SettingsNicholas Rhinehart, Rowan McAllister, Kris Kitani, Sergey LevineICCV 2019 · 407 citations
- Analyzing the Variety Loss in the Context of Probabilistic Trajectory PredictionLuca Anthony Thiede, Pratik Prabhanjan BrahmaICCV 2019 · 69 citations
- EgoEnv: Human-centric environment representations from egocentric videoTushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta Desai, James Hillis et al.NeurIPS 2023 · 28 citations
- Joint Metrics Matter: A Better Standard for Trajectory ForecastingErica Weng, Hana Hoshino, Deva Ramanan, Kris KitaniICCV 2023 · 27 citations
Builds on2
Related papers
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 69 citations
- FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video ObservationsChristian Diller, Thomas A. Funkhouser, Angela DaiCVPR 2024
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 28 citations
- Forecasting Characteristic 3D Poses of Human ActionsChristian Diller, Thomas A. Funkhouser, Angela DaiCVPR 2022 · 22 citations
- Learning State-Aware Visual Representations from Audible InteractionsHimangi Mittal, Pedro Morgado, Unnat Jain, Abhinav GuptaNeurIPS 2022 · 30 citations
