Seeing Conversations: Communication Context Identification in Egocentric Video
Tobias Dorszewski, Jens Hjortkjær
Abstract
In everyday conversations, humans effortlessly recognize communication partners using visual cues such as gaze or head orientation. Replicating this social reasoning in computer vision is challenging, especially in dynamic, multi-person settings. We introduce Communication Context Identification (CCI) in egocentric vision: Given a firstperson video sequence, determine which individuals are engaged in communication with the camera wearer. To support CCI, we collected a challenging large-scale dataset comprising 68.9 hours of egocentric video captured across diverse multi-person, multi-conversation scenarios. We propose CoCoNet, a temporal interaction model for CCI that tracks social dynamics via attention across individuals over long time scales. CoCoNet flexibly handles varying group sizes, maintains predictions through occlusions, and performs robustly even with limited temporal input. Leveraging long temporal contexts, it achieves 96% balanced accuracy on CCI. Performance varies with group size and spatial scene layout, highlighting the importance of dataset diversity. Our work advances vision-based conversational awareness, enabling applications in assistive hearing that use egocentric video to enhance individuals in the user's conversation group.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7ca9e43-a954-4d31-8aff-4b2649c1d850Builds on6
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Improving Social Awareness Through DANTE: Deep Affinity Network for Clustering Conversational InteractantsMason Swofford, John Peruzzi, Nathan Tsoi, Sydney Thompson et al.CSCW 2020 · 41 citations
- Egocentric Auditory Attention Localization in ConversationsFiona Ryan, Hao Jiang, Abhinav Shukla, James M. Rehg et al.CVPR 2023
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric PerspectiveWenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla et al.CVPR 2024
- Detecting Attended Visual Targets in VideoEunji Chong, Yongxin Wang, Nataniel Ruiz, James M. RehgCVPR 2020
Related papers
- Understanding Human Gaze Communication by Spatio-Temporal Graph ReasoningLifeng Fan, Wenguan Wang, Song-Chun Zhu, Xinyu Tang et al.ICCV 2019 · 124 citations
- EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object UnderstandingChenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei et al.ICCV 2023 · 45 citations
- Friends-MMC: A Dataset for Multi-modal Multi-party Conversation UnderstandingYueqian Wang, Xiaojun Meng, Yuxuan Wang, Jianxin Liang et al.AAAI 2025 · 6 citations
- Multi-Task Gaze Communication UnderstandingCheng Peng, Oya ÇeliktutanACM MM 2025
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard et al.NeurIPS 2024 · 24 citations
