Cross-Modal Guided Visual Synthesis for Data-Efficient Multimodal Depression Recognition
Shanliang Yang, Xiaoxiao Wang
Abstract
The performance of multimodal learning systems, particularly in high-stakes domains like automated depression recognition, is fundamentally constrained by the challenge of learning robust visual representations from limited and complex clinical data. To overcome this, we introduce Cross-Modal Guided Visual Synthesis (CMG-VS), a novel training framework that internally enhances the learning process by synthesizing new, task-relevant visual features. At its core, CMG-VS leverages the rich context from audio and text modalities to guide a conditional generative model. This model learns the intricate mapping from speech and language to visual expression, generating a diverse manifold of plausible visual behaviors to enrich the training distribution. Crucially, this synthesis is not a separate pre-processing step. Through a task-guided joint optimization scheme, the generative process is dynamically steered by the downstream multimodal recognizer's performance. This closed-loop feedback mechanism ensures the synthesized visual features are optimized to be maximally discriminative for the recognition task, rather than merely realistic. Comprehensive experiments on the widely-used DAIC-WOZ and E-DAIC benchmark datasets demonstrate that CMG-VS significantly outperforms existing state-of-the-art methods across all standard regression and classification metrics. Ablation studies further validate that our task-guided synthesis is the key driver of this performance gain, proving its effectiveness as a new paradigm for robust multimodal representation learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81689330-4c75-4c65-88db-c5cb1cc8df15Builds on7
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression EstimationHao Sun, Hongyi Wang, Jiaqing Liu, Yen-Wei Chen et al.ACM MM 2022 · 158 citations
- FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial LandmarksRuiqi Wang, Jinyang Huang, Jie Zhang, Xin Liu et al.ACM MM 2024 · 22 citations
- Diffusemix: Label-Preserving Data Augmentation with Diffusion ModelsKhawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik NandakumarCVPR 2024
Related papers
- Boosting Masked ECG-Text Auto-Encoders as Discriminative LearnersHung Manh Pham, Aaqib Saeed, Dong MaICML 2025
- Unveiling the Landscape of Clinical Depression Assessment: From Behavioral Signatures to Psychiatric ReasoningZhuang Chen, Guanqun Bi, Wen Zhang, Jiawei Hu et al.AAAI 2026 · 2 citations
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go et al.EMNLP 2025
- When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression DetectionXiangyu Zhang, Hexin Liu, Kaishuai Xu, Qiquan Zhang et al.EMNLP 2024 · 13 citations
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation LearningLuoyi Sun, Xuenan Xu, Mengyue Wu, Weidi XieACM MM 2024 · 23 citations
