Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer Fusion
Yikai Wang, Fuchun Sun, Ming Lu, Anbang Yao
摘要
We propose a compact and effective framework to fuse multimodal features at multiple layers in a single network. The framework consists of two innovative fusion schemes. Firstly, unlike existing multimodal methods that necessitate individual encoders for different modalities, we verify that multimodal features can be learnt within a shared single network by merely maintaining modality-specific batch normalization layers in the encoder, which also enables implicit fusion via joint feature representation learning. Secondly, we propose a bidirectional multi-layer fusion scheme, where multimodal features can be exploited progressively. To take advantage of such scheme, we introduce two asymmetric fusion operations including channel shuffle and pixel shift, which learn different fused features with respect to different fusion directions. These two operations are parameter-free and strengthen the multimodal feature interactions across channels as well as enhance the spatial feature discrimination within channels. We conduct extensive experiments on semantic segmentation and image translation tasks, based on three publicly available datasets covering diverse modalities. Results indicate that our proposed framework is general, compact and is superior to state-of-the-art fusion frameworks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu 等NeurIPS 2020 · 被引用 321 次
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang 等CVPR 2022 · 被引用 214 次
- DIFNet: Boosting Visual Information Flow for Image CaptioningMingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou 等CVPR 2022 · 被引用 68 次
- Rethinking Reverse Distillation for Multi-Modal Anomaly DetectionZhihao Gu, Jiangning Zhang, Liang Liu, Xu Chen 等AAAI 2024 · 被引用 51 次
- Prompting Multi-Modal Image Segmentation with Semantic GroupingQibin HeAAAI 2024 · 被引用 21 次
它引用的顶会 Paper1
相关 Paper
- StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic SegmentationBingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao 等ACM MM 2025 · 被引用 15 次
- Multi-Modal Object Re-identification via Sparse Mixture-of-ExpertsYingying Feng, Jie Li, Chi Xie, Lei Tan 等ICML 2025
- ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic SegmentationQiang Zhang, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang 等CVPR 2021
- BSAFusion: A Bidirectional Stepwise Feature Alignment Network for Unaligned Medical Image FusionHuafeng Li, Dayong Su, Qing Cai, Yafei ZhangAAAI 2025 · 被引用 43 次
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu 等ICML 2024 · 被引用 64 次
