Cross-Linked Unified Embedding for cross-modality representation learning
Xinming Tu, Zhi-Jie Cao, Chenrui Xia, Sara Mostafavi, Ge Gao
Abstract
Multi-modal learning is essential for understanding information in the real world. Jointly learning from multi-modal data enables global integration of both shared and modality-specific information, but current strategies often fail when observations from certain modalities are incomplete or missing for part of the subjects. To learn comprehensive representations based on such modality-incomplete data, we present a semi-supervised neural network model called CLUE (Cross-Linked Unified Embedding). Extending from multi-modal VAEs, CLUE introduces the use of cross-encoders to construct latent representations from modality-incomplete observations. Representation learning for modality-incomplete observations is common in genomics. For example, human cells are tightly regulated across multiple related but distinct modalities such as DNA, RNA, and protein, jointly defining a cell’s function. We benchmark CLUE on multi-modal data from single cell measurements, illustrating CLUE’s superior performance in all assessed categories of the NeurIPS 2021 Multimodal Single-cell Data Integration Competition. While we focus on analysis of single cell genomic datasets, we note that the proposed cross-linked embedding strategy could be readily applied to other cross-modality representation learning problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67d31a30-dbbb-454a-b90b-268affb5047aCited by top-tier papers5
- FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent SpaceShengzhong Liu, Tomoyoshi Kimura, Dongxin Liu, Ruijie Wang et al.NeurIPS 2023 · 72 citations
- Comprehensive View Embedding Learning for Single-Cell Multimodal IntegrationZhenchao Tang, Jiehui Huang, Guanxing Chen, Calvin Yu-Chian ChenAAAI 2024 · 11 citations
- Learning Disentangled Representation for Multi-Modal Time-Series Sensing SignalsRuichu Cai, Zhifan Jiang, Kaitao Zheng, Zijian Li et al.WWW 2025 · 8 citations
- Semi-supervised Knowledge Transfer Across Multi-omic Single-cell DataFan Zhang, Tianyu Liu, Zihao Chen, Xiaojiang Peng et al.NeurIPS 2024 · 7 citations
- Learning Crossmodal Interaction Patterns via Attributed Bipartite Graphs for Single-Cell OmicsXiaotang Wang, Xuanwei Lin, Yun Zhu, Hao Li et al.NeurIPS 2025 · 1 citation
Builds on3
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng et al.ACM MM 2020 · 93 citations
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang et al.CVPR 2020
Related papers
- Learning Multimodal VAEs through Mutual SupervisionTom Joy, Yuge Shi, Philip H. S. Torr, Tom Rainforth et al.ICLR 2022 · 27 citations
- scSSL-Bench: Benchmarking Self-Supervised Learning for Single-Cell DataOlga Ovcharenko, Florian Barkmann, Philip Toma, Imant Daunhawer et al.ICML 2025
- scMRDR: A scalable and flexible framework for unpaired single-cell multi-omics data integrationJianle Sun, Chaoqi Liang, Ran Wei, Peng Zheng et al.NeurIPS 2025 · 5 citations
- sciLaMA: A Single-Cell Representation Learning Framework to Leverage Prior Knowledge from Large Language ModelsHongru Hu, Shuwen Zhang, Yongin Choi, Venkat S. Malladi et al.ICML 2025
- MuSe-GNN: Learning Unified Gene Representation From Multimodal Biological Graph DataTianyu Liu, Yuge Wang, Rex Ying, Hongyu ZhaoNeurIPS 2023 · 44 citations
