Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal Space
Sicheng Zhao, Yaxian Li, Xingxu Yao, Weizhi Nie, Pengfei Xu, Jufeng Yang, Kurt Keutzer
摘要
Both images and music can convey rich semantics and are widely used to induce specific emotions. Matching images and music with similar emotions might help to make emotion perceptions more vivid and stronger. Existing emotion-based image and music matching methods either employ limited categorical emotion states which cannot well reflect the complexity and subtlety of emotions, or train the matching model using an impractical multi-stage pipeline. In this paper, we study end-to-end matching between image and music based on emotions in the continuous valence-arousal (VA) space. First, we construct a large-scale dataset, termed Image-Music-Emotion-Matching-Net (IMEMNet), with over 140K image-music pairs. Second, we propose cross-modal deep continuous metric learning (CDCML) to learn a shared latent embedding space which preserves the cross-modal similarity relationship in the continuous matching space. Finally, we refine the embedding space by further preserving the single-modal emotion relationship in the VA spaces of both images and music. The metric learning in the embedding space and task regression in the label space are jointly optimized for both cross-modal matching and single-modal VA prediction. The extensive experiments conducted on IMEMNet demonstrate the superiority of CDCML for emotion-based image and music matching as compared to the state-of-the-art approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation LearningJiayun Hu, Yueyi He, Tianyi Liang, Changbo Wang 等ACM MM 2025 · 被引用 2 次
- Harmonic Canvas: Inversion-Free Editing for Visually-Guided Music Style TransferYue Lei, Siqi Yang, Ting Zhong, Fan ZhouCVPR 2026
- MPJudge: Towards Perceptual Assessment of Music-Induced PaintingsShiqi Jiang, Tianyi Liang, Huayuan Ye, Changbo Wang 等AAAI 2026
它引用的顶会 Paper3
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated VideosSicheng Zhao, Yunsheng Ma, Yang Gu, Jufeng Yang 等AAAI 2020 · 被引用 123 次
- Zero-Shot Emotion Recognition via Affective Structural EmbeddingChi Zhan, Dongyu She, Sicheng Zhao, Ming-Ming Cheng 等ICCV 2019 · 被引用 54 次
- Attention-Aware Polarity Sensitive Embedding for Affective Image RetrievalXingxu Yao, Dongyu She, Sicheng Zhao, Jie Liang 等ICCV 2019 · 被引用 31 次
相关 Paper
- A Unimodal Valence-Arousal Driven Contrastive Learning Framework for Multimodal Multi-Label Emotion RecognitionWenjie Zheng, Jianfei Yu, Rui XiaACM MM 2024 · 被引用 8 次
- MusER: Musical Element-Based Regularization for Generating Symbolic Music with EmotionShulei Ji, Xinyu YangAAAI 2024 · 被引用 9 次
- Learning Visual Emotion Representations From Web DataZijun Wei, Jianming Zhang, Zhe Lin, Joon-Young Lee 等CVPR 2020
- EMOD: A Unified EEG Emotion Representation Framework Leveraging V-A Guided Contrastive LearningYuning Chen, Sha Zhao, Shijian Li, Gang PanAAAI 2026
- Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-LearningDengming Zhang, Weitao You, Ziheng Liu, Lingyun Sun 等AAAI 2025 · 被引用 2 次
