POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-view World
Boshen Xu, Sipeng Zheng, Qin Jin
Abstract
We humans are good at translating third-person observations of hand-object interactions (HOI) into an egocentric view. However, current methods struggle to replicate this ability of view adaptation from third-person to first-person. Although some approaches attempt to learn view-agnostic representation from large-scale video datasets, they ignore the relationships among multiple third-person views. To this end, we propose a Prompt-Oriented View-agnostic learning (POV) framework in this paper, which enables this view adaptation with few egocentric videos. Specifically, We introduce interactive masking prompts at the frame level to capture finegrained action information, and view-aware prompts at the token level to learn view-agnostic representation. To verify our method, we establish two benchmarks for transferring from multiple thirdperson views to the egocentric view. Our extensive experiments on these benchmarks demonstrate the efficiency and effectiveness of our POV framework and prompt tuning techniques in terms of view adaptation and view generalization. Our code is available at https://github.com/xuboshen/pov_acmmm2023.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- EgoDTM: Towards 3D-Aware Egocentric Video-Language PretrainingBoshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng et al.NeurIPS 2025 · 6 citations
- EgoPrompt: Prompt Learning for Egocentric Action RecognitionHuaihai Lyu, Chaofan Chen, Yuheng Ji, Changsheng XuACM MM 2025 · 3 citations
- Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?Boshen Xu, Ziheng Wang, Yang Du, Zhinan Song et al.ICLR 2025
Builds on26
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video RepresentationsJungin Park, Jiyoung Lee, Kwanghoon SohnCVPR 2025
- Precise Action-to-Video Generation Through Visual Action PromptsYuang Wang, Chao Wen, Haoyu Guo, Sida Peng et al.ICCV 2025 · 1 citation
- Is Tracking Really More Challenging in First Person Egocentric Vision?Matteo Dunnhofer, Zaira Manigrasso, Christian MicheloniICCV 2025 · 1 citation
- EgoPCA: A New Framework for Egocentric Hand-Object Interaction UnderstandingYue Xu, Yong-Lu Li, Zhemin Huang, Michael Xu Liu et al.ICCV 2023 · 15 citations
- Retrieval-Augmented Egocentric Video CaptioningJilan Xu, Yifei Huang, Junlin Hou, Guo Chen et al.CVPR 2024 · 16 citations
