Pose as a Modality: A Psychology-Inspired Network for Personality Recognition with a New Multimodal Dataset
Bin Tang, Keqi Pan, Miao Zheng, Ning Zhou, Jialu Sui, Dandan Zhu, Cheng-Long Deng, Shu-Guang Kuai
摘要
In recent years, predicting Big Five personality traits from multimodal data has received significant attention in artificial intelligence (AI). However, existing computational models often fail to achieve satisfactory performance. Psychological research has shown a strong correlation between pose and personality traits, yet previous research has largely ignored pose data in computational models. To address this gap, we develop a novel multimodal dataset that incorporates full-body pose data. The dataset includes video recordings of 287 participants completing a virtual interview with 36 questions, along with self-reported Big Five personality scores as labels. To effectively utilize this multimodal data, we introduce the Psychology-Inspired Network (PINet), which consists of three key modules: Multimodal Feature Awareness (MFA), Multimodal Feature Interaction (MFI), and Psychology-Informed Modality Correlation Loss (PIMC Loss). The MFA module leverages the Vision Mamba Block to capture comprehensive visual features related to personality, while the MFI module efficiently fuses the multimodal features. The PIMC Loss, grounded in psychological theory, guides the model to emphasize different modalities for different personality dimensions. Experimental results show that the PINet outperforms several state-of-the-art baseline models. Furthermore, the three modules of PINet contribute almost equally to the model’s overall performance. Incorporating pose data significantly enhances the model’s performance, with the pose modality ranking mid-level in importance among the five modalities. These findings address the existing gap in personality-related datasets that lack full-body pose data and provide a new approach for improving the accuracy of personality prediction models, highlighting the importance of integrating psychological insights into AI frameworks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang 等NeurIPS 2021 · 被引用 782 次
- Low-Rank Bottleneck in Multi-head Attention ModelsSrinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi 等ICML 2020 · 被引用 130 次
- Investigating how speech and animation realism influence the perceived personality of virtual characters and agentsSean Thomas, Ylva Ferstl, Rachel McDonnell, Cathy EnnisIEEE VR 2022 · 被引用 34 次
相关 Paper
- Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level ScoresRyo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka 等AAAI 2025 · 被引用 3 次
- Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT ImagesJie Mei, Chenyu Lin, Yu Qiu, Yaonan Wang 等CVPR 2025
- PersonalitySensing: A Multi-View Multi-Task Learning Approach for Personality Detection based on Smartphone UsageSongcheng Gao, Wenzhong Li, Lynda J. Song, Xiao Zhang 等ACM MM 2020 · 被引用 14 次
- PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment AnalysisHeng Xie, Kang Zhu, Zhengqi Wen, Jianhua Tao 等AAAI 2026 · 被引用 1 次
- Robust Pose Estimation in Crowded Scenes with Direct Pose-Level InferenceDongkai Wang, Shiliang Zhang, Gang HuaNeurIPS 2021 · 被引用 36 次
