Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer Fusion
Yikai Wang, Fuchun Sun, Ming Lu, Anbang Yao
Abstract
We propose a compact and effective framework to fuse multimodal features at multiple layers in a single network. The framework consists of two innovative fusion schemes. Firstly, unlike existing multimodal methods that necessitate individual encoders for different modalities, we verify that multimodal features can be learnt within a shared single network by merely maintaining modality-specific batch normalization layers in the encoder, which also enables implicit fusion via joint feature representation learning. Secondly, we propose a bidirectional multi-layer fusion scheme, where multimodal features can be exploited progressively. To take advantage of such scheme, we introduce two asymmetric fusion operations including channel shuffle and pixel shift, which learn different fused features with respect to different fusion directions. These two operations are parameter-free and strengthen the multimodal feature interactions across channels as well as enhance the spatial feature discrimination within channels. We conduct extensive experiments on semantic segmentation and image translation tasks, based on three publicly available datasets covering diverse modalities. Results indicate that our proposed framework is general, compact and is superior to state-of-the-art fusion frameworks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 272d798c-ff79-4ad2-9347-227edbf17a19Cited by top-tier papers13
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu et al.NeurIPS 2020 · 321 citations
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang et al.CVPR 2022 · 214 citations
- DIFNet: Boosting Visual Information Flow for Image CaptioningMingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou et al.CVPR 2022 · 68 citations
- Rethinking Reverse Distillation for Multi-Modal Anomaly DetectionZhihao Gu, Jiangning Zhang, Liang Liu, Xu Chen et al.AAAI 2024 · 51 citations
- Prompting Multi-Modal Image Segmentation with Semantic GroupingQibin HeAAAI 2024 · 21 citations
Builds on1
Related papers
- StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic SegmentationBingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao et al.ACM MM 2025 · 15 citations
- Multi-Modal Object Re-identification via Sparse Mixture-of-ExpertsYingying Feng, Jie Li, Chi Xie, Lei Tan et al.ICML 2025
- ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic SegmentationQiang Zhang, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang et al.CVPR 2021
- BSAFusion: A Bidirectional Stepwise Feature Alignment Network for Unaligned Medical Image FusionHuafeng Li, Dayong Su, Qing Cai, Yafei ZhangAAAI 2025 · 43 citations
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu et al.ICML 2024 · 64 citations
