A Progressive Skip Reasoning Fusion Method for Multi-Modal Classification
Qian Guo, Xinyan Liang, Yuhua Qian, Zhihua Cui, Jie Wen
Abstract
In multi-modal classification tasks, a good fusion algorithm can effectively integrate and process multi-modal data, thereby significantly improving its performance. Researchers often focus on the design of complex fusion operators and have proposed numerous fusion operators, while paying less attention to the design of feature fusion usage, specifically how features should be fused to better facilitate multi-modal classification tasks. In this article, we propose a progressive skip reasoning fusion network (PSRFN) to make some attempts to address this issue. Firstly, unlike most existing multi-modal fusion methods that only use one fusion operator in a single stage to fuse all view features, PSRFN utilizes the progressive skip reasoning (PSR) block to fuse all views with a fusion operator at each layer. Specifically, each PSR block utilizes all view features and the fused features from the previous layer to jointly obtain the fused features for the current layer. Secondly, each PSR block utilizes a dual-weighted fusion strategy with learnable parameters to adaptively allocate weights during the fusion process. The first level of weighting assigns weights to each view feature, while the second level assigns weights to the fused features from the previous layer and the fused features obtained from the first level of weighting in the current layer. This strategy ensures that the PSR block can dynamically adjust the weights based on the actual contribution of features. Finally, to enable the model to fully utilize feature information from different levels for feature fusion, the skip connections are adopted between PSR blocks. Extensive experiment results on six real multi-modal datasets show that a better usage for fusion operator is indeed able to improve performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6670dcd8-8323-4999-ac40-e5604710338cCited by top-tier papers6
- Uncertainty-Guided View-Strength-Aware Feature Utilization for Multi-View ClassificationLi Lv, Qian Guo, Li Zhang, Liang Du et al.AAAI 2026
- Robust Signal Enhancement via Fractional Detail Views and Knowledge Guided Multi-view FusionZikun Jin, Yuhua Qian, Xinyan Liang, Jiaqian Zhang et al.ICML 2026
- Multi-View Clustering with Granularity-Aware Pseudo SupervisionJie Yang, Cheng-You Lu, Zhongli Wang, Hsiang-Ting Chen et al.AAAI 2026
- Evolutionary Multi-View Classification with Label Noise via Gradient and Feature Dual-PerceptionShuai Li, Xinyan Liang, Yuhua Qian, Li LvICML 2026
- Incomplete Multi-View Clustering via Neighborhood-Conditioned DiffusionQian Guo, Gaohui Zuo, Bingbing Jiang, Guangrui Fan et al.ICML 2026
Related papers
- SMR-Net: Semantic-Guided Mutually Reinforcing Network for Cross-Modal Image Fusion and Salient Object DetectionGuobao Xiao, Xinyu Liu, Zebin Lin, Rui MingAAAI 2025 · 11 citations
- Deep Embedded Complementary and Interactive Information for Multi-View ClassificationJinglin Xu, Wenbin Li, Xinwang Liu, Dingwen Zhang et al.AAAI 2020 · 65 citations
- Knowledge-Enhanced Multimodal Fake News Detection: Semantic Visual and Priority FusionQin Zhang, Jiaying Liu, Qian Tao, Zhiwei Guo et al.WWW 2026
- PFFN: Progressive Feature Fusion Network for Lightweight Image Super-ResolutionDongyang Zhang, Changyu Li, Ning Xie, Guoqing Wang et al.ACM MM 2021 · 18 citations
- Progressive Modality Reinforcement for Human Multimodal Emotion Recognition From Unaligned Multimodal SequencesFengmao Lv, Xiang Chen, Yanyong Huang, Lixin Duan et al.CVPR 2021
