VIBE: Disentangling Social Dynamics via Kinematics-Informed Variational Inference for Behavioral Emotion
Abhishek Pratap Singh, Vaibhav Singh, Deepak Kumar, Balasubramanian Raman
摘要
Group Emotion Recognition (GER) is crucial for understanding social dynamics, ranging from interpreting intimate conversations to evaluating crowd behavior in large-scale surveillance scenarios. While current AI models can analyze these scenes, they often act as black boxes that take shortcuts. Instead of focusing on how people are actually behaving, these models often get distracted by the background environment, leading to inaccurate results. To bridge this gap, we introduce VIBE (Variational Inference for Behavioral Emotion), a kinematics-aware framework that integrates audio, video, and text modalities through causal structuring. Unlike standard models that simply mix data together, VIBE utilizes mathematical constraints to filter out background noise and isolate the genuine emotions of the people involved. This purified representation enables our model to focus exclusively on the sociological mechanics of the crowd, dynamically modulating neural attention based on raw physical synchrony. Simultaneously, we align visual dynamics with human interpretability by projecting latent representations into a semantically structured space informed by textual descriptions. Comprehensive experiments demonstrate that VIBE consistently outperforms state-of-the-art methods. Code is available at GitHub.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du 等ACM MM 2022 · 被引用 260 次
- Everything at Once - Multi-modal Fusion Transformer for Video RetrievalNina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas 等CVPR 2022 · 被引用 4 次
- VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingLimin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong 等CVPR 2023
相关 Paper
- ViBE: Dressing for Diverse Body ShapesWei-Lin Hsiao, Kristen GraumanCVPR 2020
- Grounding Emotion Recognition with Visual Prototypes: VEGA - Revisiting CLIP in MERCGuanyu Hu, Dimitrios Kollias, Xinyu YangACM MM 2025 · 被引用 5 次
- Bring the VibeOn: Designing a Multimodal Interface for Shared Emotional Experiences in Live-streamed ConcertsGyeongjin Kim, Sebin Lee, Daye Kim, Jungjin Lee 等ACM MM 2025 · 被引用 2 次
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu 等ACM MM 2025 · 被引用 9 次
- Unsupervised Learning From Video With Deep Neural EmbeddingsChengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark 等CVPR 2020
