Advancing Radiograph Representation Learning with Masked Record Modeling
Hong-Yu Zhou, Chenyu Lian, Liansheng Wang, Yizhou Yu
摘要
Modern studies in radiograph representation learning (R 2 L) rely on either selfsupervision to encode invariant semantics or associated radiology reports to incorporate medical expertise, while the complementarity between them is barely noticed. To explore this, we formulate the self-and report-completion as two complementary objectives and present a unified framework based on masked record modeling (MRM). In practice, MRM reconstructs masked image patches and masked report tokens following a multi-task scheme to learn knowledge-enhanced semantic representations. With MRM pre-training, we obtain pre-trained models that can be well transferred to various radiography tasks. Specifically, we find that MRM offers superior performance in label-efficient fine-tuning. For instance, MRM achieves 88.5% mean AUC on CheXpert using 1% labeled data, outperforming previous R 2 L methods with 100% labels. On NIH ChestX-ray, MRM outperforms the best performing counterpart by about 3% under small labeling ratios. Besides, MRM surpasses self-and report-supervised pre-training in identifying the pneumonia type and the pneumothorax area, sometimes by large margins. Code and models are available at https://github.com/RL4M/ MRM-pytorch .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing BiasZhongwei Wan, Che Liu, Mi Zhang, Jie Fu 等NeurIPS 2023 · 被引用 114 次
- G2D: From Global to Dense Radiography Representation Learning via Vision-Language Pre-trainingChe Liu, Cheng Ouyang, Sibo Cheng, Anand Shah 等NeurIPS 2024 · 被引用 21 次
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-Language Pre-TrainingWeiwei Cao, Jianpeng Zhang, Zhongyi Shui, Sinuo Wang 等ICCV 2025 · 被引用 18 次
- Unlocking the Power of Spatial and Temporal Information in Medical Multimodal Pre-trainingJinxia Yang, Bing Su, Xin Zhao, Ji-Rong WenICML 2024 · 被引用 13 次
- Boosting Medical Visual Understanding From Multi-Granular Language LearningZihan Li, Yiqing Wang, Sina Farsiu, Paul KinahanICLR 2026 · 被引用 6 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui 等ICLR 2022 · 被引用 565 次
- GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image RecognitionShih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena YeungICCV 2021 · 被引用 516 次
相关 Paper
- CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus DatasetXiao Wang, Fuling Wang, Yuehang Li, Qingchuan Ma 等CVPR 2025
- MRM: Masked Relation Modeling for Medical Image Pre-Training with GeneticsQiushi Yang, Wuyang Li, Baopu Li, Yixuan YuanICCV 2023 · 被引用 20 次
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report GenerationPuzhen Wu, Hexin Dong, Yi Lin, Yihao Ding 等AAAI 2026 · 被引用 3 次
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report GenerationZhuoxiao Chen, Hongyang Yu, Ying Xu, Yadan Luo 等CVPR 2026 · 被引用 3 次
- RadLAS: A Foundation Model for Interpretable Radiography Image Analysis with Lesion-Aware Self-Supervised Pre-trainingYihang Liu, Ying Wen, Longzhen Yang, Lianghua He 等ACM MM 2025
