Learning from Label Relationships in Human Affect
Niki Maria Foteinopoulou, Ioannis Patras
Abstract
Human affect and mental state estimation in an automated manner, face a number of difficulties, including learning from labels with poor or no temporal resolution, learning from few datasets with little data (often due to confidentiality constraints) and, (very) long, in-the-wild videos. For these reasons, deep learning methodologies tend to overfit, that is, arrive at latent representations with poor generalisation performance on the final regression task. To overcome this, in this work, we introduce two complementary contributions. First, we introduce a novel relational loss for multilabel regression and ordinal problems that regularises learning and leads to better generalisation. The proposed loss uses label vector inter-relational information to learn better latent representations by aligning batch label distances to the distances in the latent feature space. Second, we utilise a two-stage attention architecture that estimates a target for each clip by using features from the neighbouring clips as temporal context. We evaluate the proposed methodology on both continuous affect and schizophrenia severity estimation problems, as there are methodological and contextual parallels between the two. Experimental results demonstrate that the proposed methodology outperforms the baselines that are trained using the supervised regression loss, as well as pre-training the network architecture with an unsupervised contrastive loss. In the domain of schizophrenia, the proposed methodology outperforms previous state-of-the-art by a large margin, achieving a PCC of up to 78% performance close to that of human experts (85%) and much higher than previous works (uplift of up to 40%). In the case of affect recognition, we outperform previous vision-based methods in terms of CCC on both the OMG and the AMIGOS datasets. Specifically for AMIGOS, we outperform previous SoTA CCC for both arousal and valence by 9% and 13% respectively, and in the OMG dataset we outperform previous vision works by up to 5% for both arousal and valence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa2be915-d3a4-4e30-adab-9643629fd7d4Cited by top-tier papers1
Ask how each one uses itBuilds on8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Multi-Label Classification with Label Graph SuperimposingYa Wang, Dongliang He, Fu Li, Xiang Long et al.AAAI 2020 · 192 citations
- MIMAMO Net: Integrating Micro- and Macro-Motion for Video Emotion RecognitionDidan Deng, Zhaokang Chen, Yuqian Zhou, Bertram E. ShiAAAI 2020 · 51 citations
Related papers
- MDDR: Multi-modal Dual-Attention aggregation for Depression RecognitionWei Zhang, En Zhu, Juan Chen, Yunpeng LiACM MM 2024 · 8 citations
- MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution DistillationZhicheng Zhang, Pancheng Zhao, Eunil Park, Jufeng YangCVPR 2024 · 11 citations
- Dep-MAP: A Multi-level Alignment Framework with Semantic Prototypes for Video-based Automatic Depression AssessmentHao Wang, Jiayu Ye, Qingxiang WangAAAI 2026
- DHCM-CACL: Dynamic Hierarchical Cross-modal Mamba with Confidence-Adaptive Contrastive Learning for Multimodal Emotion RecognitionBaiqiang Wu, Yang LiAAAI 2026 · 1 citation
- Representation Learning through Multimodal Attention and Time-Sync Comments for Affective Video Content AnalysisJicai Pan, Shangfei Wang, Lin FangACM MM 2022 · 18 citations
