Tuning Medical Foundation Models for Inner Ear Temporal CT Analysis with Plug-and-play Domain Knowledge Aggregator
Weixun Wan, Xinyang Jiang, Zilong Wang, Bei Li, Cairong Zhao
Abstract
High-resolution computed tomography (CT) is essential for diagnosing hearing loss and planning interventions such as cochlear implantation, as it provides detailed visualization of inner-ear anatomy. This paper focuses on advancing AI-based analysis of inner-ear CT scans to support clinical decision-making. However, a major challenge lies in the scarcity of annotated data, which limits the applicability of conventional supervised learning techniques. To address this, we present the first publicly available Children's Inner Ear CT Dataset (CIED), comprising 722 CT scans labeled for structural anomaly detection, postoperative hearing outcome prediction, and anatomical segmentation. In addition, we explore the use of medical foundation models to improve generalization in data-scarce scenarios. Existing parameter-efficient adaptation methods often fall short in two ways: they lack a unified mechanism to adapt across diverse foundation model architectures and they are not specifically designed to incorporate domain expert knowledge of inner-ear anatomy and pathology. To overcome these limitations, we propose Domain Knowledge Guided Tuning (DKGT), a plug-and-play framework that introduces a unified adapter—Domain Knowledge Aggregator (DKA)—to inject radiomics-based anatomical features into foundation models via cross-attention. DKA supports various backbone types and preserves pretrained representations of foundation model while enabling multi-layer integration of expert knowledge. Extensive experiments across multiple tasks demonstrate that DKGT consistently outperforms state-of-the-art classification methods, achieving superior performance and generalizability on inner-ear CT analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He et al.ICLR 2023 · 204 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
- Lightweight Vision Transformer with Bidirectional InteractionQihang Fan, Huaibo Huang, Xiaoqiang Zhou, Ran HeNeurIPS 2023 · 59 citations
- How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?Wenxuan Li, Alan L. Yuille, Zongwei ZhouICLR 2024 · 21 citations
Related papers
- Bridging Radiology and Pathology Foundation Models via Concept-Based Multimodal Co-AdaptationYihang Chen, Yanyan Huang, Fuying Wang, Maximus Yeung et al.ICLR 2026
- PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric ApplicationsDingkang Yang, Jinjie Wei, Dongling Xiao, Shunli Wang et al.NeurIPS 2024 · 40 citations
- Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-SupervisionYunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu et al.CVPR 2026
- DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical ImagingYuxi Ma, Sujie Liu, Jing Yang, Jiacheng Wang et al.CVPR 2026
- SegVol: Universal and Interactive Volumetric Medical Image SegmentationYuxin Du, Fan Bai, Tiejun Huang, Bo ZhaoNeurIPS 2024 · 155 citations
