Lune

ACM MM2025顶会

Modal Symbiosis: Variational Alignment Unveils New Horizons in Multimodal Representation Learning

Zeyan Li, Cankun Guo, Yin Tang

2025年份

摘要

Multimodal models integrate visual, textual, and other data to achieve human-like understanding, but this fusion creates a conflict between cross-modal alignment and modality-specific expertise.The pursuit of unified feature spaces often undermines specialized knowledge in individual modalities, as shown by performance drops in unimodal tasks. To resolve this contradiction, we propose VAMP (Variational Alignment with Modality Preservation), a novel multimodal framework featuring a Dynamic Feature Diversion mechanism that partitions modal representations into two components-one preserving modality-specific expertise and the other enabling cross-modal alignment. Inspired by Variational Canonical Correlation Analysis, we introduce a shared space projection layer that maps features into a common representational space while preserving modality-specific characteristics. We further implement a Progressive Training Strategy that sequentially freezes different components before full fine-tuning, preventing mode collapse and enhancing generalization capabilities. Experimental results demonstrate VAMP's significant performance improvements across zero-shot image classification, cross-modal retrieval, and visual question answering, while simultaneously outperforming baseline models on unimodal tasks. This research provides an engineered solution to the ''knowledge dilution'' problem in cross-modal alignment.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖