RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification
Shakhrul Iman Siam, Tiantian Feng, Jiankun Zhang, Shrikanth Narayanan, Mi Zhang
摘要
Respiratory diseases remain a leading cause of global mortality, where timely and accurate diagnosis is critical to improving patient outcomes and reducing healthcare burdens. While prior work has explored audio-based models for respiratory disease detection, such unimodal approaches often suffer from limited generalizability and diagnostic precision. In this paper, we propose RespiraMFM, a Multimodal Foundation Model that integrates respiratory sounds with patient medical history and symptoms to enhance diagnostic accuracy and disease detection capabilities. We introduce an effective contrastive alignment strategy for audio-text multimodal integration, allowing the model to learn better cross-modal representations between respiratory sounds and corresponding textual clinical information. We evaluate RespiraMFM across five major respiratory diseases using seven real-world datasets in both supervised fine-tuning and zero-shot settings, achieving a 9.15% improvement in AUROC on supervised tasks and a 20.98% gain on zero-shot tasks over existing baselines. These findings underscore the potential of our framework to advance early diagnosis and improve clinical decision-making in respiratory disease management.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Soundwave: Less is More for Speech-Text Alignment in LLMsYuhao Zhang, Zhiheng Liu, Fan Bu, Ruiyu Zhang 等ACL 2025 · 被引用 11 次
- UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them AllYuanhuiyi Lyu, Xu Zheng, Jiazhou Zhou, Lin WangCVPR 2024 · 被引用 8 次
相关 Paper
- Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease DiagnosisPengfei ZHANG, Tianxin Xie, Minghao Yang, Li LiuICLR 2026
- TAMER: A Tri-Modal Contrastive Alignment and Multi-Scale Embedding Refinement Framework for Zero-Shot ECG DiagnosisXuewei Zhou, Yajie Meng, Pan Zeng, Xianfang Tang 等CVPR 2026
- FLAM: Frame-Wise Language-Audio ModelingYusong Wu, Christos Tsirigotis, Ke Chen, Cheng-Zhi Anna Huang 等ICML 2025
- ProbMED: A Probabilistic Framework for Medical Multimodal BindingYuan Gao, Sangwook Kim, Jianzhong You, Chris McIntoshICCV 2025 · 被引用 3 次
- SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory SignalsRahul Thapa, Bryan He, Magnus Ruud Kjær, Hyatt E. Moore IV 等ICML 2024 · 被引用 48 次
