DFCNet: Dual-Factor Compensatory Clustering Network for Modality-Imbalanced Generalized Zero-Shot Learning
Xiangyu Shan, Heng Song, Junwu Zhu
摘要
In audio-visual joint analysis, generalized zero-shot learning (GZSL) aims to recognize unseen categories by aligning semantic information across modalities. However, significant challenges arise from temporal and semantic discrepancies between audio and video modalities. Traditional methods typically depend on static semantic embeddings, thereby overlooking the dynamic nature of these modalities and often resulting in modality imbalance. We propose the Dual-Factor Compensatory Clustering Network (DFCNet), an end-to-end framework for dynamic fusion and optimized alignment of heterogeneous modal information to address these limitations. DFCNet employs a multi-branch architecture, integrating a parallel multi-layer perceptron (MLP) for semantic modeling and a Bidirectional Long Short-Term Memory (BiLSTM) network for capturing temporal consistency. The Compensatory Fusion Block (CFB) employs tensor decomposition to facilitate cross-modal coupling, where the low-rank representation decomposition aligns intra-modal feature distributions. Additionally, we introduce the Dual-Factor Clustering Multi-objective Optimization Framework (DCMOF), which ensures gradient equilibrium and adaptively adjusts modality contribution weights to strengthen robust cross-modal alignment. Designed as a pluggable module, DFCNet can be seamlessly integrated into existing base models. Experimental results demonstrate that our framework significantly improves the performance of the Audio-Visual Cross-Modal Alignment (AVCA) model across multiple GZSL datasets. Ablation studies further validate the critical role of CFB in cross-modal alignment and highlight the significance of DCMOF in optimizing modality coordination. The code is available at https://github.com/ATKEROM/DFCNet.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot LearningChaoqun Wang, Xuejin Chen, Shaobo Min, Xiaoyan Sun 等AAAI 2021 · 被引用 22 次
- Discrepancy-Aware Attention Network for Enhanced Audio-Visual Generalized Zero-Shot LearningRunlin Yu, Yipu Gong, Wenrui Li, Aiwen Sun 等ACM MM 2025 · 被引用 1 次
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 被引用 43 次
- Audiovisual Generalised Zero-shot Learning with Cross-modal Attention and LanguageOtniel-Bogdan Mercea, Lukas Riesch, A. Sophia Koepke, Zeynep AkataCVPR 2022 · 被引用 54 次
- Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot LearningMan Liu, Feng Li, Chunjie Zhang, Yunchao Wei 等CVPR 2023
