POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-view World
Boshen Xu, Sipeng Zheng, Qin Jin
摘要
We humans are good at translating third-person observations of hand-object interactions (HOI) into an egocentric view. However, current methods struggle to replicate this ability of view adaptation from third-person to first-person. Although some approaches attempt to learn view-agnostic representation from large-scale video datasets, they ignore the relationships among multiple third-person views. To this end, we propose a Prompt-Oriented View-agnostic learning (POV) framework in this paper, which enables this view adaptation with few egocentric videos. Specifically, We introduce interactive masking prompts at the frame level to capture finegrained action information, and view-aware prompts at the token level to learn view-agnostic representation. To verify our method, we establish two benchmarks for transferring from multiple thirdperson views to the egocentric view. Our extensive experiments on these benchmarks demonstrate the efficiency and effectiveness of our POV framework and prompt tuning techniques in terms of view adaptation and view generalization. Our code is available at https://github.com/xuboshen/pov_acmmm2023.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- EgoDTM: Towards 3D-Aware Egocentric Video-Language PretrainingBoshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng 等NeurIPS 2025 · 被引用 6 次
- EgoPrompt: Prompt Learning for Egocentric Action RecognitionHuaihai Lyu, Chaofan Chen, Yuheng Ji, Changsheng XuACM MM 2025 · 被引用 3 次
- Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?Boshen Xu, Ziheng Wang, Yang Du, Zhinan Song 等ICLR 2025
它引用的顶会 Paper26
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
相关 Paper
- Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video RepresentationsJungin Park, Jiyoung Lee, Kwanghoon SohnCVPR 2025
- Precise Action-to-Video Generation Through Visual Action PromptsYuang Wang, Chao Wen, Haoyu Guo, Sida Peng 等ICCV 2025 · 被引用 1 次
- Is Tracking Really More Challenging in First Person Egocentric Vision?Matteo Dunnhofer, Zaira Manigrasso, Christian MicheloniICCV 2025 · 被引用 1 次
- EgoPCA: A New Framework for Egocentric Hand-Object Interaction UnderstandingYue Xu, Yong-Lu Li, Zhemin Huang, Michael Xu Liu 等ICCV 2023 · 被引用 15 次
- Retrieval-Augmented Egocentric Video CaptioningJilan Xu, Yifei Huang, Junlin Hou, Guo Chen 等CVPR 2024 · 被引用 16 次
