Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching Network
Kai Cheng, Xin Liu, Yiu-ming Cheung, Rui Wang, Xing Xu, Bineng Zhong
Abstract
Many cognitive researches have shown that human may 'see voices' or 'hear faces', and such ability can be potentially associated by machine vision and intelligence. However, this research is still under early stage. In this paper, we present a novel adversarial deep semantic matching network for efficient voice-face interactions and associations, which can well learn the correspondence between voices and faces for various cross-modal matching and retrieval tasks. Within the proposed framework, we exploit a simple and efficient adversarial learning architecture to learn the cross-modal embeddings between faces and voices, which consists of two subnetworks, respectively, for generator and discriminator. The former subnetwork is designed to adaptively discriminate the high-level semantical features between voices and faces, in which the triplet loss and multi-modal center loss are in tandem utilized to explicitly regularize the correspondences among them. The latter subnetwork is further leveraged to maximally bridge the semantic gap between the representations of voice and face data, featuring on maintaining the semantic consistency. Through the joint exploitation of the above, the proposed framework can well push representations of voice-face data from the same person closer while pulling those representations of different person away. Extensive experiments empirically show that the proposed approach involves fewer parameters and calculations, adapts various cross-modal matching tasks for voice-face data and brings substantial improvements over the state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Taking a Part for the Whole: An Archetype-agnostic Framework for Voice-Face AssociationGuancheng Chen, Xin Liu, Xing Xu, Yiu-Ming Cheung et al.ACM MM 2023 · 1 citation
- Seeking the Shape of Sound: An Adaptive Framework for Learning Voice-Face AssociationPeisong Wen, Qianqian Xu, Yangbangyan Jiang, Zhiyong Yang et al.CVPR 2021
- ASMR: Learning Attribute-Based Person Search with Adaptive Semantic Margin RegularizerBoseung Jeong, Jicheol Park, Suha KwakICCV 2021 · 29 citations
- From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from SpeechHyeong-Seok Choi, Changdae Park, Kyogu LeeICLR 2020 · 33 citations
- Heterogeneous Attention Network for Effective and Efficient Cross-modal RetrievalTan Yu, Yi Yang, Yi Li, Lin Liu et al.SIGIR 2021 · 50 citations
