MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
Ziyang Zhang, Yang Yu, Yucheng Chen, Xulei Yang, Si Yong Yeo
摘要
Despite significant progress in Vision-Language Pretraining (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual content. This gap hinders the model's ability to synthesize coherent and novel visual representations from textual prompts, thereby reducing the effectiveness of multi-modal learning. In this work, we propose MedUnifier, a unified VLP framework tailored for medical data. MedUnifier seamlessly integrates text-grounded image generation capabilities with multi-modal learning strategies, including image-text contrastive alignment, image-text matching and image-grounded text generation. Unlike traditional methods that reply on continuous visual representations, our approach employs visual vector quantization, which not only facilitates a more cohesive learning strategy for crossmodal understanding but also enhances multi-modal generation quality by effectively leveraging discrete representations. Our framework's effectiveness is evidenced by the experiments on established benchmarks, including uni-modal tasks, cross-modal tasks, and multi-modal tasks, where it achieves state-of-the-art performance across various tasks. MedUnifier also offers a highly adaptable tool for a wide range of language and vision tasks in healthcare, marking advancement toward the development of a generalizable AI model for medical applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 被引用 1 次
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse AttentionPeng Zhang, Zhihui Lai, Wenting Chen, Xu Wu 等AAAI 2026
- Medical Vision-Language Pretraining with LLM-Guided Temporal SupervisionLiang Bai, Zhi Wang, Huimin Yan, Xian YangAAAI 2026
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Towards Unifying Medical Vision-and-Language Pre-training via Soft PromptsZhihong Chen, Shizhe Diao, Benyou Wang, Guanbin Li 等ICCV 2023 · 被引用 50 次
- Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing BiasZhongwei Wan, Che Liu, Mi Zhang, Jie Fu 等NeurIPS 2023 · 被引用 114 次
- Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question AnsweringYuanhao Zou, Zhaozheng YinCVPR 2025
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningWei Li, Can Gao, Guocheng Niu, Xinyan Xiao 等ACL 2021
- G2D: From Global to Dense Radiography Representation Learning via Vision-Language Pre-trainingChe Liu, Cheng Ouyang, Sibo Cheng, Anand Shah 等NeurIPS 2024 · 被引用 21 次
