Human Identification and Interaction Detection in Cross-View Multi-Person Videos with Wearable Cameras
Jiewen Zhao, Ruize Han, Yiyang Gan, Liang Wan, Wei Feng, Song Wang
Abstract
Compared to a single fixed camera, multiple moving cameras, e.g., those worn by people, can better capture the human interactive and group activities in a scene, by providing multiple, flexible and possibly complementary views of the involved people. In this setting the actual promotion of activity detection is highly dependent on the effective correlation and collaborative analysis of multiple videos taken by different wearable cameras, which is highly challenging given the time-varying view differences across different cameras and mutual occlusion of people in each video. By focusing on two wearable cameras and the interactive activities that involve only two people, in this paper we develop a new approach that can simultaneously: (i) identify the same persons across the two videos, (ii) detect the interactive activities of interest, including their occurrence intervals and involved people, and (iii) recognize the category of each interactive activity. Specifically, we represent each video by a graph, with detected persons as nodes, and propose a unified Graph Neural Network (GNN) based framework to jointly solve the above three problems. A graph matching network is developed for identifying the same persons across the two videos and a graph inference network is then used for detecting the human interactions. We also build a new video dataset, which provides a benchmark for this study, and conduct extensive experiments to validate the effectiveness and superiority of the proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 34d67661-1349-4c1f-b313-e6fb16768451Cited by top-tier papers3
- Multiview Detection with Shadow Transformer (and View-Coherent Data Augmentation)Yunzhong Hou, Liang ZhengACM MM 2021 · 65 citations
- Self-supervised Multi-view Multi-Human Association and TrackingYiyang Gan, Ruize Han, Liqiang Yin, Wei Feng et al.ACM MM 2021 · 44 citations
- Connecting the Complementary-view Videos: Joint Camera Identification and Subject AssociationRuize Han, Yiyang Gan, Jiacheng Li, Feifan Wang et al.CVPR 2022 · 12 citations
Related papers
- Complementary-View Co-Interest Person DetectionRuize Han, Jiewen Zhao, Wei Feng, Yiyang Gan et al.ACM MM 2020 · 18 citations
- Fusing Personal and Environmental Cues for Identification and Segmentation of First-Person Camera Wearers in Third-Person ViewsZiwei Zhao, Yuchen Wang, Chuhua WangCVPR 2024
- Complementary-View Multiple Human TrackingRuize Han, Wei Feng, Jiewen Zhao, Zicheng Niu et al.AAAI 2020 · 36 citations
- Towards a Dynamic Inter-Sensor Correlations Learning Framework for Multi-Sensor-Based Wearable Human Activity RecognitionShenghuan Miao, Ling Chen, Rong Hu, Yingsong LuoUbiComp 2022 · 35 citations
- Self-Supervised Human Pose based Multi-Camera Video SynchronizationLiqiang Yin, Ruize Han, Wei Feng, Song WangACM MM 2022 · 7 citations
