F-Assist: Multi-Phase Fetal Growth Forecast and Report Generation from Ultrasound Examination
Bin Pu, Xusheng Liang, Xinpeng Ding, Jinlin Wu, Zhen Lei, Shengli Li, Kenli Li, Jiawei Ma
Abstract
Forecasting fetal growth from sequential ultrasound examinations is essential for personalized prenatal care. Existing medical vision-language models (MLLMs) are limited to single-phase/organ evaluations and qualitative reasoning, neglecting longitudinal history and precise continuous biometric values. To address this gap, we introduce the novel task that integrates multi-modal ultrasound examination results at early pregnance stage to predict the subsequent fetal health status. To support this task, we first present GrowthFetus, the interesting multi-phase, multi-organ fetal ultrasound dataset to date, containing 9,280 examinations from 2,000 fetuses. Based on this dataset, we propose F 2 -Assist, a unified MLLM framework with three key components: (i) a Cross-Phase Organ Alignment module for heterogeneous multi-organ feature fusion across phases, (ii) a History-Aware Temporal Encoding module for modeling irregular temporal dynamics, and (iii) a Growth Parameter Adapter that encodes continuous biometric values as differentiable tokens for numerically precise reasoning. Extensive experiments show that F 2 -Assist achieves temporally coherent predictions and clinically consistent reports, significantly outperforming state-of-the-art MLLMs. Our study establishes a practical framework for longitudinal ultrasound analysis, bridging growth forecasting and report generation in a unified model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 552 citations
- LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and GenerationSuhyeon Lee, Won Jun Kim, Jinho Chang, Jong Chul YeICLR 2024 · 80 citations
Related papers
- Organ-Aware Routing Mixture-of-Retrieval Augmented Generation for Fetal Ultrasound ReportingBin Pu, Siyu Wang, Rongbin Li, Xinpeng Ding et al.AAAI 2026
- FETAL-GAUGE: A BENCHMARK FOR ASSESSING VISION-LANGUAGE MODELS IN FETAL ULTRASOUNDHussain Alasmawi, Numan Saeed, Mohammad YaqubICLR 2026
- EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound IntelligenceChaoyin She, Ruifang Lu, Lida Chen, Wei Wang et al.ACL 2026 · 8 citations
- LLaVA-Ultra: Large Chinese Language and Vision Assistant for UltrasoundXuechen Guo, Wenhao Chai, Shiyan Li, Gaoang WangACM MM 2024 · 18 citations
- U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound UnderstandingAnjie Le, Henan Liu, Yue Wang, Zhenyu Liu et al.ICLR 2026 · 8 citations
