Best of Both Worlds: Multimodal Contrastive Learning with Tabular and Imaging Data
Paul Hager, Martin J. Menten, Daniel Rueckert
Abstract
Medical datasets and especially biobanks, often contain extensive tabular data with rich clinical information in addition to images. In practice, clinicians typically have less data, both in terms of diversity and scale, but still wish to deploy deep learning solutions. Combined with increasing medical dataset sizes and expensive annotation costs, the necessity for unsupervised methods that can pretrain multimodally and predict unimodally has risen.
To address these needs, we propose the first selfsupervised contrastive learning framework that takes advantage of images and tabular data to train unimodal encoders. Our solution combines SimCLR and SCARF, two leading contrastive learning strategies, and is simple and effective. In our experiments, we demonstrate the strength of our framework by predicting risks of myocardial infarction and coronary artery disease (CAD) using cardiac MR images and 120 clinical features from 40,000 UK Biobank subjects. Furthermore, we show the generalizability of our approach to natural images using the DVM car advertisement dataset.
We take advantage of the high interpretability of tabular data and through attribution and ablation experiments find that morphometric tabular features, describing size and shape, have outsized importance during the contrastive learning process and improve the quality of the learned embeddings. Finally, we introduce a novel form of supervised contrastive learning, label as a feature (LaaF), by appending the ground truth label as a tabular feature during multimodal pretraining, outperforming all supervised contrastive baselines. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b87bbd91-deee-4e61-8b01-6c20dfb34d78Cited by top-tier papers13
- SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-trainingKazem Meidani, Parshin Shojaee, Chandan K. Reddy, Amir Barati FarimaniICLR 2024 · 37 citations
- Tabular Insights, Visual Impacts: Transferring Expertise from Tables to ImagesJun-Peng Jiang, Han-Jia Ye, Leye Wang, Yang Yang et al.ICML 2024 · 15 citations
- RegBN: Batch Normalization of Multimodal Data with RegularizationMorteza Ghahremani, Christian WachingerNeurIPS 2023 · 15 citations
- MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular LearningWall Kim, Chaeyoung Song, Hanul KimCVPR 2026 · 9 citations
- Hierarchical Pretraining on Multimodal Electronic Health RecordsXiaochen Wang, Junyu Luo, Jiaqi Wang, Ziyi Yin et al.EMNLP 2023 · 7 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
Related papers
- No Data? No Problem: Robust Vision-Tabular Learning with Missing ValuesMarta Hasny, Laura Daza, Keno Bressem, Maxime Di Folco et al.ICML 2026
- Scarf: Self-Supervised Contrastive Learning using Random Feature CorruptionDara Bahri, Heinrich Jiang, Yi Tay, Donald MetzlerICLR 2022 · 233 citations
- Sequential Multi-Dimensional Self-Supervised Learning for Clinical Time SeriesAniruddh Raghu, Payal Chandak, Ridwan Alam, John V. Guttag et al.ICML 2023 · 18 citations
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 907 citations
- ContIG: Self-supervised Multimodal Contrastive Learning for Medical Imaging with GeneticsAiham Taleb, Matthias Kirchler, Remo Monti, Christoph LippertCVPR 2022 · 64 citations
