Bimodal masked language modeling for bulk RNA-seq and DNA methylation representation learning
Maxence Gélard, Hakim Benkirane, Thomas Pierrot, Guillaume Richard, Paul-Henry Cournède
摘要
Oncologists are increasingly relying on multiple modalities to model the complexity of diseases. Within this landscape, transcriptomic and epigenetic data have proven to be particularly instrumental and play an increasingly vital role in clinical applications. However, their integration into multimodal models remains a challenge, especially considering their high dimensionality. In this work, we present a novel bimodal model that jointly learns representations of bulk RNA-seq and DNA methylation leveraging self-supervision from masked language modeling. We implement an architecture that reduces the memory footprint usually attributed to purely transformer-based models when dealing with long sequences. We demonstrate that the obtained bimodal embeddings can be used to fine-tune cancer-type classification and survival models that achieve state-of-the-art performance compared to unimodal models. Furthermore, we introduce a robust learning framework that maintains downstream task performance despite missing modalities, enhancing the model’s applicability in real-world clinical settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov 等AAAI 2021 · 被引用 393 次
- Are Multimodal Transformers Robust to Missing Modality?Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine 等CVPR 2022 · 被引用 153 次
- Multi-modal Transfer Learning between Biological Foundation ModelsJuan Jose Garau-Luis, Patrick Bordes, Liam Gonzalez, Masa Roller 等NeurIPS 2024 · 被引用 19 次
- Test-Time Adaptation for Combating Missing Modalities in Egocentric VideosMerey Ramazanova, Alejandro Pardo, Bernard Ghanem, Motasem AlfarraICLR 2025
相关 Paper
- Modaltune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-Task Learning in Digital PathologyVishwesh Ramanathan, Tony Xu, Pushpak Pati, Faruk Ahmed 等ICCV 2025 · 被引用 4 次
- MUST: Modality-Specific Representation-Aware Transformer for Diffusion-Enhanced Survival Prediction with Missing ModalityKyungwon Kim, Dosik HwangCVPR 2026 · 被引用 1 次
- HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataKonstantin Hemker, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 被引用 80 次
- A New Paradigm for Genome-wide DNA Methylation Prediction Without Methylation InputXiaoke Huang, Qi Liu, Yifei Zhao, Xianfeng Tang 等ICLR 2026 · 被引用 2 次
- Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival PredictionGuillaume Jaume, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson 等CVPR 2024
