Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease Diagnosis
Yuxin Lin, Haoran Li, Haoyu Cao, Yongting Hu, Qihao Xu, Chengliang Liu, Xiaoling Luo, Zhihao Wu, Yong Xu, Wei Wang
Abstract
Multimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention mechanisms - treat cross-modal interactions indiscriminately, lacking adaptive modulation of modality-specific contributions under varying clinical scenarios. We propose an adaptive fusion framework that dynamically routes and refines multimodal signals for enhancing disease recognition. The framework comprises two key components: 1) Dynamic Cross-Modal Expert Routing (CMER), which selectively activates convolutional neural network (CNN) experts from one modality based on contextual guidance from the other, ensuring only the most relevant feature extractors contribute to fusion; and 2) Top-K Expert-Guided Wavelet Fusion (TEWF), which performs discrete wavelet transform (DWT) to decompose selected features into low- and high-frequency subbands. Cross-modal attention is then applied specifically to high-frequency components, where lesion-specific microstructures reside, enabling frequency-aware fusion. Finally, inverse DWT (IDWT) reconstructs the fused representation, weighted by CMER-derived importance scores to amplify informative modality cues while suppressing redundancy. Experimental validation on two multimodal retinal datasets demonstrates that our method achieves state-of-the-art performance, outperforming existing fusion strategies by significant margins in disease classification accuracy and robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc5fa60d-3fc6-4f78-9416-0420a270670eBuilds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 110 citations
- MVCINN: Multi-View Diabetic Retinopathy Detection Using a Deep Cross-Interaction Neural NetworkXiaoling Luo, Chengliang Liu, Waikeung Wong, Jie Wen et al.AAAI 2023 · 14 citations
- Scalable One-Pass Incomplete Multi-View Clustering by Aligning AnchorsYalan Qin, Guorui Feng, Xinpeng ZhangAAAI 2025 · 3 citations
- Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy GradingXiaoling Luo, Qihao Xu, Huisi Wu, Chengliang Liu et al.AAAI 2025 · 3 citations
Related papers
- Multi-Modal Multi-Instance Learning for Retinal Disease RecognitionXirong Li, Yang Zhou, Jie Wang, Hailan Lin et al.ACM MM 2021 · 52 citations
- Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus DiseasesYuxin Lin, Wei Wang, Xiaoling Luo, Zhihao Wu et al.AAAI 2025 · 3 citations
- OmniFM: Toward Modality-Robust and Task-Agnostic Federated Learning for Heterogeneous Medical ImagingMeilin Liu, Jiaying Wang, Jing ShanCVPR 2026 · 1 citation
- CA-MLIF: Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework for Tumor Prognostic PredictionYajun An, Jiale Chen, Huan Lin, Zhenbing Liu et al.AAAI 2025 · 1 citation
- Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image SegmentationFan Zhang, Zhiwei Gu, Hua WangAAAI 2026 · 4 citations
