Seeking the Shape of Sound: An Adaptive Framework for Learning Voice-Face Association
Peisong Wen, Qianqian Xu, Yangbangyan Jiang, Zhiyong Yang, Yuan He, Qingming Huang
摘要
Nowadays, we have witnessed the early progress on learning the association between voice and face automatically, which brings a new wave of studies to the computer vision community. However, most of the prior arts along this line (a) merely adopt local information to perform modality alignment and (b) ignore the diversity of learning difficulty across different subjects. In this paper, we propose a novel framework to jointly address the above-mentioned issues. Targeting at (a), we propose a two-level modality alignment loss where both global and local information are considered. Compared with the existing methods, we introduce a global loss into the modality alignment process. The global component of the loss is driven by the identity classification. Theoretically, we show that minimizing the loss could maximize the distance between embeddings across different identities while minimizing the distance between embeddings belonging to the same identity, in a global sense (instead of a mini-batch). Targeting at (b), we propose a dynamic reweighting scheme to better explore the hard but valuable identities while filtering out the unlearnable identities. Experiments show that the proposed method outperforms the previous methods in multiple settings, including voice-face matching, verification and retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object DetectionRunmin Cong, Hongyu Liu, Chen Zhang, Wei Zhang 等ACM MM 2023 · 被引用 71 次
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng 等ACM MM 2022 · 被引用 34 次
- Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?Cho-Ying Wu, Chin-Cheng Hsu, Ulrich NeumannCVPR 2022 · 被引用 16 次
- Regularized Contrastive Partial Multi-view Outlier DetectionYijia Wang, Qianqian Xu, Yangbangyan Jiang, Siran Dai 等ACM MM 2024 · 被引用 8 次
- Rethinking Voice-Face Correlation: A Geometry ViewXiang Li, Yandong Wen, Muqiao Yang, Jinglu Wang 等ACM MM 2023 · 被引用 4 次
它引用的顶会 Paper3
- From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from SpeechHyeong-Seok Choi, Changdae Park, Kyogu LeeICLR 2020 · 被引用 33 次
- Speech Fusion to Face: Bridging the Gap Between Human's Vocal Characteristics and Facial ImagingYeqi Bai, Tao Ma, Lipo Wang, Zhenjie ZhangACM MM 2022 · 被引用 12 次
- Circle Loss: A Unified Perspective of Pair Similarity OptimizationYifan Sun, Changmao Cheng, Yuhan Zhang, Chi Zhang 等CVPR 2020
相关 Paper
- Taking a Part for the Whole: An Archetype-agnostic Framework for Voice-Face AssociationGuancheng Chen, Xin Liu, Xing Xu, Yiu-Ming Cheung 等ACM MM 2023 · 被引用 1 次
- Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching NetworkKai Cheng, Xin Liu, Yiu-ming Cheung, Rui Wang 等ACM MM 2020 · 被引用 19 次
- Learning Concordant Attention via Target-aware Alignment for Visible-Infrared Person Re-identificationJianbing Wu, Hong Liu, Yuxin Su, Wei Shi 等ICCV 2023 · 被引用 45 次
- Spatial-Frequency Collaborative Learning for Occluded Visible-Infrared Person Re-IdentificationJIan Yu, Yujian Feng, Shuai You, Zhongkai Zhou 等CVPR 2026
- Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image MatchingYafei Zhang, Yongle Shang, Huafeng LiACM MM 2025 · 被引用 5 次
