Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease Diagnosis
Yuxin Lin, Haoran Li, Haoyu Cao, Yongting Hu, Qihao Xu, Chengliang Liu, Xiaoling Luo, Zhihao Wu, Yong Xu, Wei Wang
摘要
Multimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention mechanisms - treat cross-modal interactions indiscriminately, lacking adaptive modulation of modality-specific contributions under varying clinical scenarios. We propose an adaptive fusion framework that dynamically routes and refines multimodal signals for enhancing disease recognition. The framework comprises two key components: 1) Dynamic Cross-Modal Expert Routing (CMER), which selectively activates convolutional neural network (CNN) experts from one modality based on contextual guidance from the other, ensuring only the most relevant feature extractors contribute to fusion; and 2) Top-K Expert-Guided Wavelet Fusion (TEWF), which performs discrete wavelet transform (DWT) to decompose selected features into low- and high-frequency subbands. Cross-modal attention is then applied specifically to high-frequency components, where lesion-specific microstructures reside, enabling frequency-aware fusion. Finally, inverse DWT (IDWT) reconstructs the fused representation, weighted by CMER-derived importance scores to amplify informative modality cues while suppressing redundancy. Experimental validation on two multimodal retinal datasets demonstrates that our method achieves state-of-the-art performance, outperforming existing fusion strategies by significant margins in disease classification accuracy and robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 被引用 110 次
- MVCINN: Multi-View Diabetic Retinopathy Detection Using a Deep Cross-Interaction Neural NetworkXiaoling Luo, Chengliang Liu, Waikeung Wong, Jie Wen 等AAAI 2023 · 被引用 14 次
- Scalable One-Pass Incomplete Multi-View Clustering by Aligning AnchorsYalan Qin, Guorui Feng, Xinpeng ZhangAAAI 2025 · 被引用 3 次
- Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy GradingXiaoling Luo, Qihao Xu, Huisi Wu, Chengliang Liu 等AAAI 2025 · 被引用 3 次
相关 Paper
- Multi-Modal Multi-Instance Learning for Retinal Disease RecognitionXirong Li, Yang Zhou, Jie Wang, Hailan Lin 等ACM MM 2021 · 被引用 52 次
- Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus DiseasesYuxin Lin, Wei Wang, Xiaoling Luo, Zhihao Wu 等AAAI 2025 · 被引用 3 次
- OmniFM: Toward Modality-Robust and Task-Agnostic Federated Learning for Heterogeneous Medical ImagingMeilin Liu, Jiaying Wang, Jing ShanCVPR 2026 · 被引用 1 次
- CA-MLIF: Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework for Tumor Prognostic PredictionYajun An, Jiale Chen, Huan Lin, Zhenbing Liu 等AAAI 2025 · 被引用 1 次
- Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image SegmentationFan Zhang, Zhiwei Gu, Hua WangAAAI 2026 · 被引用 4 次
