FETAL-GAUGE: A BENCHMARK FOR ASSESSING VISION-LANGUAGE MODELS IN FETAL ULTRASOUND
Hussain Alasmawi, Numan Saeed, Mohammad Yaqub
Abstract
The growing demand for prenatal ultrasound imaging has intensified a global shortage of trained sonographers, creating barriers to essential fetal health monitoring. Deep learning has the potential to enhance sonographers' efficiency and support the training of new practitioners. Vision-Language Models (VLMs) are particularly promising for ultrasound interpretation, as they can jointly process images and text to perform multiple clinical tasks within a single framework. However, despite the expansion of VLMs, no standardized benchmark exists to evaluate their performance in fetal ultrasound imaging. This gap is primarily due to the modality’s challenging nature, operator dependency, and the limited public availability of datasets. To address this gap, we present Fetal-Gauge, the first and largest visual question answering benchmark specifically designed to evaluate VLMs across various fetal ultrasound tasks. Our benchmark comprises over 42,000 images and 93,000 question-answer pairs, spanning anatomical plane identification, visual grounding of anatomical structures, fetal orientation assessment, clinical view conformity, and clinical diagnosis. We systematically evaluate several state-of-the-art VLMs, including general-purpose and medical-specific models, and reveal a substantial performance gap: the best-performing model achieves only 55% accuracy, far below clinical requirements. Our analysis identifies critical limitations of current VLMs in fetal ultrasound interpretation, highlighting the urgent need for domain-adapted architectures and specialized training approaches. Fetal-Gauge establishes a rigorous foundation for advancing multimodal deep learning in prenatal care and provides a pathway toward addressing global healthcare accessibility challenges. Our benchmark is publicly available at https://github.com/BioMedIA-MBZUAI/FETAL-GAUGE
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29b8bb0a-38e0-4a26-8dac-fd3c9a4a6ed8Builds on3
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual PromptsMu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer et al.CVPR 2024
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMYutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao et al.CVPR 2024
Related papers
- U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound UnderstandingAnjie Le, Henan Liu, Yue Wang, Zhenyu Liu et al.ICLR 2026 · 8 citations
- EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound IntelligenceChaoyin She, Ruifang Lu, Lida Chen, Wei Wang et al.ACL 2026 · 8 citations
- F-Assist: Multi-Phase Fetal Growth Forecast and Report Generation from Ultrasound ExaminationBin Pu, Xusheng Liang, Xinpeng Ding, Jinlin Wu et al.CVPR 2026
- Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive EvaluationHong-Tao Yu, Yuxin Peng, Serge J. Belongie, Xiu-Shen WeiICLR 2026 · 21 citations
- OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical TasksZhihao Peng, Cheng Wang, Shengyuan Liu, Zhiying Liang et al.CVPR 2026 · 7 citations
