PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
Danyal Maqbool, Changhee Lee, Zachary Huemann, Samuel Church, Matthew E. Larson, Scott B. Perlman, Tomas A. Romero, Joshua D. Warner, Meghan G. Lubner, Xin Tie, Jameson Merkow, Junjie Hu
Abstract
Generating automated reports for 3D positron emission tomography (PET) is an important and challenging task in medical imaging. PET plays a vital role in oncology, but automating report generation is difficult due to the complexity of whole-body 3D volumes, the wide range of potential clinical findings, and the limited availability of annotated datasets. To address these challenges, we introduce PETARSeg-11K, the first large-scale, publicly available dataset that provides lesion-level correspondence between 3D PET/CT volumes and free-text radiological findings. It comprises 11,356 lesion descriptions paired with 3D segmentations. Second, we propose PETAR-4B, a 3D vision-language model designed for mask-aware, spatially grounded PET/CT reporting. PETAR-4B jointly encodes PET, CT, and 3D lesion segmentation masks, using a 3D focal prompt to capture fine-grained details of lesions that normally comprise less than 0.1% of the volume. Evaluations using automated metrics show PETAR-4B substantially outperforming all 2D and 3D baselines. A human study involving five physicians -- the first of its kind for automated PET reporting -- confirms the model's clinical utility and establishes correlations between automated metrics and expert judgment. This work provides a foundational dataset and a novel architecture, advancing 3D medical vision-language understanding in PET.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO TrainingDavid Dai, Peilin Chen, Chanakya Ekbote, Paul Pu LiangNeurIPS 2025 · 48 citations
Related papers
- PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission TomographyYichi Zhang, Wenbo Zhang, Zehui Ling, Gang Feng et al.AAAI 2026 · 4 citations
- VoxTell: Free-Text Promptable Universal 3D Medical Image SegmentationMaximilian Rokuss, Moritz Langenberg, Yannick Kirchhoff, Fabian Isensee et al.CVPR 2026 · 22 citations
- SegAnyPET: Universal Promptable Segmentation from Positron Emission Tomography ImagesYichi Zhang, Le Xue, Wenbo Zhang, Lanlan Li et al.ICCV 2025 · 7 citations
- Anatomical Region-Guided 3D PET/MR Tumor Segmentation via Medical RecordTianming Xu, Tiantian Guo, Youdan Feng, Zihan Chen et al.ACM MM 2025
- RadGPT: Constructing 3D Image-Text Tumor DatasetsPedro R. A. S. Bassi, Mehmet Can Yavuz, Ibrahim Ethem Hamamci, Sezgin Er et al.ICCV 2025 · 48 citations
