Fusing Personal and Environmental Cues for Identification and Segmentation of First-Person Camera Wearers in Third-Person Views
Ziwei Zhao, Yuchen Wang, Chuhua Wang
Abstract
As wearable cameras become more popular, an important question emerges: how to identify camera wearers within the perspective of conventional static cameras. The drastic difference between first-person (egocentric) and third-person (exocentric) camera views makes this a challenging task. We present PersonEnvironmentNet (PEN), a framework designed to integrate information from both the individuals in the two views and geometric cues inferred from the background environment. To facilitate research in this direction, we also present TF2023, a novel dataset comprising synchronized first-person and third-person views, along with masks of camera wearers and labels associating these masks with the respective first-person views. In addition, we propose a novel quantitative metric designed to measure a model's ability to comprehend the relationship between the two views. Our experiments reveal that PEN outperforms existing methods. The code and dataset are available at https://github.com/ ziweizhao1993/PEN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1f97487-333a-449a-adff-478122af8115Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 191 citations
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He et al.CVPR 2022 · 168 citations
Related papers
- Human Identification and Interaction Detection in Cross-View Multi-Person Videos with Wearable CamerasJiewen Zhao, Ruize Han, Yiyang Gan, Liang Wan et al.ACM MM 2020 · 32 citations
- EgoHumans: An Egocentric 3D Multi-Human BenchmarkRawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe et al.ICCV 2023 · 59 citations
- Estimating Egocentric 3D Human Pose in the Wild with External Weak SupervisionJian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar et al.CVPR 2022 · 33 citations
- Egocentric Pose Estimation from Human Vision SpanHao Jiang, Vamsi Krishna IthapuICCV 2021 · 37 citations
- EgoEnv: Human-centric environment representations from egocentric videoTushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta Desai, James Hillis et al.NeurIPS 2023 · 28 citations
