DPNET: Dynamic Poly-attention Network for Trustworthy Multi-modal Classification
Xin Zou, Chang Tang, Xiao Zheng, Zhenglai Li, Xiao He, Shan An, Xinwang Liu
Abstract
With advances in sensing technology, multi-modal data collected from different sources are increasingly available. Multi-modal classification aims to integrate complementary information from multi-modal data to improve model classification performance. However, existing multi-modal classification methods are basically weak in integrating global structural information and providing trustworthy multi-modal fusion, especially in safety-sensitive practical applications (e.g., medical diagnosis). In this paper, we propose a novel Dynamic Poly-attention Network (DPNET) for trustworthy multi-modal classification. Specifically, DPNET has four merits: (i) To capture the intrinsic modality-specific structural information, we design a structure-aware feature aggregation module to learn the corresponding structure-preserved global compact feature representation. (ii) A transparent fusion strategy based on the modality confidence estimation strategy is induced to track information variation within different modalities for dynamical fusion. (iii) To facilitate more effective and efficient multi-modal fusion, we introduce a cross-modal low-rank fusion module to reduce the complexity of tensor-based fusion and activate the implication of different rank-wise features via a rank attention mechanism. (iv) A label confidence estimation module is devised to drive the network to generate more credible confidence. An intra-class attention loss is introduced to supervise the network training. Extensive experiments on four real-world multi-modal biomedical datasets demonstrate that the proposed method achieves competitive performance compared to other state-of-the-art ones.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers8
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context RetentionXin Zou, Di Lu, Yizhou Wang, Yibo Yan et al.NeurIPS 2025 · 49 citations
- Enhancing Multi-View Classification Reliability with Adaptive RejectionWei Liu, Yufei Chen, Xiaodong YueAAAI 2025 · 6 citations
- SparseMVC: Probing Cross-view Sparsity Variations for Multi-view ClusteringRuimeng Liu, Xin Zou, Chang Tang, Xiao Zheng et al.NeurIPS 2025 · 5 citations
- Who Should I Trust? Explicit Confidence-Focused Multimodal Intent RecognitionYi Liu, Qimeng Yang, Lanlan LuAAAI 2026
- A Peer-review Look on Multi-modal Clustering: An Information Bottleneck Realization MethodZhengzheng Lou, Hang Xue, Chaoyang Zhang, Shizhe HuICML 2025
Related papers
- Multi-Level Confidence Learning for Trustworthy Multimodal ClassificationXiao Zheng, Chang Tang, Zhiguo Wan, Chengyu Hu et al.AAAI 2023 · 41 citations
- Multimodal Dynamics: Dynamical Fusion for Trustworthy Multimodal ClassificationZongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang et al.CVPR 2022 · 149 citations
- CL-DMDF: Dynamic Multimodal Data Fusion Model Based on Contrastive LearningDong Li, Lingling Zhang, Binghao Han, Linlin Ding et al.AAAI 2026
- Scalable Medical Multimodal Fusion via Symmetric Consistency ModelingXiaowen Sun, Hui Liu, Gongguan Chen, Ning MaoICML 2026
- HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataKonstantin Hemker, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 80 citations
