Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
Qi Chen, Xinze Zhou, Chen Liu, Hao Chen, Wenxuan Li, Zekun Jiang, Ziyan Huang, Yuxuan Zhao, Dexin Yu, Junjun He, Yefeng Zheng, Ling Shao
摘要
AI for tumor segmentation is limited by the lack of large, voxel-wise annotated datasets, which are hard to create and require medical experts. In our proprietary JHH dataset of 3,000 annotated pancreatic tumor scans, we found that AI performance stopped improving after scans. With synthetic data, we reached the same performance using only real scans. This finding suggests that synthetic data can steepen data scaling laws, enabling more efficient model training than real data alone. Motivated by these lessons, we created AbdomenAtlas 2.0—a dataset of 10,134 CT scans with a total of 13,223 tumor instances per-voxel manually annotated in six organs (pancreas, liver, kidney, colon, esophagus, and uterus) and 6,511 control scans. Annotated by 23 expert radiologists, it is several orders of magnitude larger than existing public tumor datasets. While we continue expanding the dataset, the current version of AbdomenAtlas 2.0 already provides a strong foundation—based on lessons from the JHH dataset—for training AI to segment tumors in six organs. It achieves notable improvements over public datasets, with a DSC gain on in-distribution tests and on out-of-distribution tests.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Patient-Specific Disease Dynamics With Latent Flow Matching For Longitudinal Imaging GenerationHao Chen, Rui Yin, Yifan Chen, Qi Chen 等ICLR 2026 · 被引用 11 次
- Glance and Focus Reinforcement for Pan-cancer ScreeningLinshan Wu, Jia-Xin Zhuang, Hao ChenICLR 2026 · 被引用 2 次
- Are Pixel-Wise Metrics Reliable for Computerized Tomography Reconstruction?Tianyu Lin, Xinran Li, Chuntung Zhuang, Qi Chen 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth 等CVPR 2022 · 被引用 736 次
相关 Paper
- How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?Wenxuan Li, Alan L. Yuille, Zongwei ZhouICLR 2024 · 被引用 21 次
- Label-Free Liver Tumor SegmentationQixin Hu, Yixiong Chen, Junfei Xiao, Shuwen Sun 等CVPR 2023
- Towards Generalizable Tumor SynthesisQi Chen, Xiaoxi Chen, Haorui Song, Zhiwei Xiong 等CVPR 2024
- RadGPT: Constructing 3D Image-Text Tumor DatasetsPedro R. A. S. Bassi, Mehmet Can Yavuz, Ibrahim Ethem Hamamci, Sezgin Er 等ICCV 2025 · 被引用 48 次
- Prior-Aware Neural Network for Partially-Supervised Multi-Organ SegmentationYuyin Zhou, Zhe Li, Song Bai, Xinlei Chen 等ICCV 2019 · 被引用 196 次
