Bodily Behaviors in Social Interaction: Novel Annotations and State-of-the-Art Evaluation
Michal Balazia, Philipp Müller, Ákos Levente Tánczos, August von Liechtenstein, François Brémond
Abstract
Body language is an eye-catching social signal and its automatic analysis can significantly advance artificial intelligence systems to understand and actively participate in social interactions. While computer vision has made impressive progress in low-level tasks like head and body pose estimation, the detection of more subtle behaviors such as gesturing, grooming, or fumbling is not well explored. In this paper we present BBSI, the first set of annotations of complex Bodily Behaviors embedded in continuous Social Interactions in a group setting. Based on previous work in psychology, we manually annotated 26 hours of spontaneous human behavior in the MPIIGroupInteraction dataset with 15 distinct body language classes. We present comprehensive descriptive statistics on the resulting dataset as well as results of annotation quality evaluations. For automatic detection of these behaviors, we adapt the Pyramid Dilated Attention Network (PDAN), a state-of-the-art approach for human action detection. We perform experiments using four variants of spatial-temporal features as input to PDAN: Two-Stream Inflated 3D CNN, Temporal Segment Networks, Temporal Shift Module and Swin Transformer. Results are promising and indicate a great room for improvement in this difficult task. Representing a key piece in the puzzle towards automatic understanding of social behavior, BBSI is fully available to the research community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbf2ebc3-a8bd-4575-8af8-690bcdfe68f1Cited by top-tier papers3
- Prototypical Calibrating Ambiguous Samples for Micro-Action RecognitionKun Li, Dan Guo, Guoliang Chen, Chunxiao Fan et al.AAAI 2025 · 55 citations
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang et al.ICCV 2025 · 21 citations
- MA-Bench: Towards Fine-grained Micro-Action UnderstandingKun Li, Jihao Gu, Fei Wang, Zhiliang Wu et al.CVPR 2026 · 12 citations
Builds on7
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Toyota Smarthome: Real-World Activities of Daily LivingSrijan Das, Rui Dai, Michal Koperski, Luca Minciullo et al.ICCV 2019 · 182 citations
- 3D CNNs With Adaptive Temporal Feature ResolutionsMohsen Fayyaz, Emad Bahrami Rad, Ali Diba, Mehdi Noroozi et al.CVPR 2021
Related papers
- PANDA: A Gigapixel-Level Human-Centric Video DatasetXueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo et al.CVPR 2020
- Multi-Task Gaze Communication UnderstandingCheng Peng, Oya ÇeliktutanACM MM 2025
- What-Meets-Where: Unified Learning of Action and Contact Localization in ImagesYuxiao Wang, Yu Lei, Wolin Liang, Weiying Xue et al.AAAI 2026 · 1 citation
- SocialGesture: Delving into Multi-person Gesture UnderstandingXu Cao, Pranav Virupaksha, Wenqi Jia, Bolin Lai et al.CVPR 2025
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt et al.ICCV 2019 · 108 citations
