Seeking the Shape of Sound: An Adaptive Framework for Learning Voice-Face Association
Peisong Wen, Qianqian Xu, Yangbangyan Jiang, Zhiyong Yang, Yuan He, Qingming Huang
Abstract
Nowadays, we have witnessed the early progress on learning the association between voice and face automatically, which brings a new wave of studies to the computer vision community. However, most of the prior arts along this line (a) merely adopt local information to perform modality alignment and (b) ignore the diversity of learning difficulty across different subjects. In this paper, we propose a novel framework to jointly address the above-mentioned issues. Targeting at (a), we propose a two-level modality alignment loss where both global and local information are considered. Compared with the existing methods, we introduce a global loss into the modality alignment process. The global component of the loss is driven by the identity classification. Theoretically, we show that minimizing the loss could maximize the distance between embeddings across different identities while minimizing the distance between embeddings belonging to the same identity, in a global sense (instead of a mini-batch). Targeting at (b), we propose a dynamic reweighting scheme to better explore the hard but valuable identities while filtering out the unlearnable identities. Experiments show that the proposed method outperforms the previous methods in multiple settings, including voice-face matching, verification and retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object DetectionRunmin Cong, Hongyu Liu, Chen Zhang, Wei Zhang et al.ACM MM 2023 · 71 citations
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng et al.ACM MM 2022 · 34 citations
- Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?Cho-Ying Wu, Chin-Cheng Hsu, Ulrich NeumannCVPR 2022 · 16 citations
- Regularized Contrastive Partial Multi-view Outlier DetectionYijia Wang, Qianqian Xu, Yangbangyan Jiang, Siran Dai et al.ACM MM 2024 · 8 citations
- Rethinking Voice-Face Correlation: A Geometry ViewXiang Li, Yandong Wen, Muqiao Yang, Jinglu Wang et al.ACM MM 2023 · 4 citations
Builds on3
- From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from SpeechHyeong-Seok Choi, Changdae Park, Kyogu LeeICLR 2020 · 33 citations
- Speech Fusion to Face: Bridging the Gap Between Human's Vocal Characteristics and Facial ImagingYeqi Bai, Tao Ma, Lipo Wang, Zhenjie ZhangACM MM 2022 · 12 citations
- Circle Loss: A Unified Perspective of Pair Similarity OptimizationYifan Sun, Changmao Cheng, Yuhan Zhang, Chi Zhang et al.CVPR 2020
Related papers
- Taking a Part for the Whole: An Archetype-agnostic Framework for Voice-Face AssociationGuancheng Chen, Xin Liu, Xing Xu, Yiu-Ming Cheung et al.ACM MM 2023 · 1 citation
- Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching NetworkKai Cheng, Xin Liu, Yiu-ming Cheung, Rui Wang et al.ACM MM 2020 · 19 citations
- Learning Concordant Attention via Target-aware Alignment for Visible-Infrared Person Re-identificationJianbing Wu, Hong Liu, Yuxin Su, Wei Shi et al.ICCV 2023 · 45 citations
- Spatial-Frequency Collaborative Learning for Occluded Visible-Infrared Person Re-IdentificationJIan Yu, Yujian Feng, Shuai You, Zhongkai Zhou et al.CVPR 2026
- Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image MatchingYafei Zhang, Yongle Shang, Huafeng LiACM MM 2025 · 5 citations
