Ego-Topo: Environment Affordances From Egocentric Video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, Kristen Grauman
摘要
First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions from the persistent space itself. We introduce a model for environment affordances that is learned directly from egocentric video. The main idea is to gain a human-centric model of a physical space (such as a kitchen) that captures (1) the primary spatial zones of interaction and (2) the likely activities they support. Our approach decomposes a space into a topological map derived from first-person activity, organizing an ego-video into a series of visits to the different zones. Further, we show how to link zones across multiple related environments (e.g., from videos of multiple kitchens) to obtain a consolidated representation of environment functionality. On EPIC-Kitchens and EGTEA+, we demonstrate our approach for learning scene affordances and anticipating future actions in long-form video. Project page: http://vision.cs. utexas.edu/projects/ego-topo/ In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 被引用 270 次
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo 等ICLR 2022 · 被引用 119 次
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu 等ICLR 2024 · 被引用 93 次
- Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric ViewsVincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee 等AAAI 2021 · 被引用 89 次
- Learning Affordance Landscapes for Interaction Exploration in 3D EnvironmentsTushar Nagarajan, Kristen GraumanNeurIPS 2020 · 被引用 87 次
它引用的顶会 Paper3
- What Would You Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality AttentionAntonino Furnari, Giovanni Maria FarinellaICCV 2019 · 被引用 204 次
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 被引用 194 次
- Generative Hybrid Representations for Activity Forecasting With No-Regret LearningJiaqi Guan, Ye Yuan, Kris M. Kitani, Nicholas RhinehartCVPR 2020
相关 Paper
- EgoEnv: Human-centric environment representations from egocentric videoTushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta Desai, James Hillis 等NeurIPS 2023 · 被引用 28 次
- Multi-label affordance mapping from egocentric visionLorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-CantinICCV 2023 · 被引用 26 次
- Shaping embodied agent behavior with activity-context priors from egocentric videoTushar Nagarajan, Kristen GraumanNeurIPS 2021 · 被引用 23 次
- Human Hands as Probes for Interactive Object UnderstandingMohit Goyal, Sahil Modi, Rishabh Goyal, Saurabh GuptaCVPR 2022 · 被引用 26 次
- Ego-Exo: Transferring Visual Representations From Third-Person to First-Person VideosYanghao Li, Tushar Nagarajan, Bo Xiong, Kristen GraumanCVPR 2021
