Seeing Conversations: Communication Context Identification in Egocentric Video
Tobias Dorszewski, Jens Hjortkjær
摘要
In everyday conversations, humans effortlessly recognize communication partners using visual cues such as gaze or head orientation. Replicating this social reasoning in computer vision is challenging, especially in dynamic, multi-person settings. We introduce Communication Context Identification (CCI) in egocentric vision: Given a firstperson video sequence, determine which individuals are engaged in communication with the camera wearer. To support CCI, we collected a challenging large-scale dataset comprising 68.9 hours of egocentric video captured across diverse multi-person, multi-conversation scenarios. We propose CoCoNet, a temporal interaction model for CCI that tracks social dynamics via attention across individuals over long time scales. CoCoNet flexibly handles varying group sizes, maintains predictions through occlusions, and performs robustly even with limited temporal input. Leveraging long temporal contexts, it achieves 96% balanced accuracy on CCI. Performance varies with group size and spatial scene layout, highlighting the importance of dataset diversity. Our work advances vision-based conversational awareness, enabling applications in assistive hearing that use egocentric video to enhance individuals in the user's conversation group.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Improving Social Awareness Through DANTE: Deep Affinity Network for Clustering Conversational InteractantsMason Swofford, John Peruzzi, Nathan Tsoi, Sydney Thompson 等CSCW 2020 · 被引用 41 次
- Egocentric Auditory Attention Localization in ConversationsFiona Ryan, Hao Jiang, Abhinav Shukla, James M. Rehg 等CVPR 2023
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric PerspectiveWenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla 等CVPR 2024
- Detecting Attended Visual Targets in VideoEunji Chong, Yongxin Wang, Nataniel Ruiz, James M. RehgCVPR 2020
相关 Paper
- Understanding Human Gaze Communication by Spatio-Temporal Graph ReasoningLifeng Fan, Wenguan Wang, Song-Chun Zhu, Xinyu Tang 等ICCV 2019 · 被引用 124 次
- EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object UnderstandingChenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei 等ICCV 2023 · 被引用 45 次
- Friends-MMC: A Dataset for Multi-modal Multi-party Conversation UnderstandingYueqian Wang, Xiaojun Meng, Yuxuan Wang, Jianxin Liang 等AAAI 2025 · 被引用 6 次
- Multi-Task Gaze Communication UnderstandingCheng Peng, Oya ÇeliktutanACM MM 2025
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard 等NeurIPS 2024 · 被引用 24 次
