Lune

NeurIPS2025Top-tier venue

Scaling Language-centric Omnimodal Representation Learning

Chenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu, Mahani Aljunied, Yu Rong

2025Year
25Citations
2Top-tier citations

Abstract

Recent multimodal embedding approaches leveraging multimodal large language models (MLLMs) fine-tuned with contrastive learning (CL) have shown promising results, yet the underlying reasons behind their superiority remain underexplored. This work argues that a crucial advantage of MLLM-based approaches stems from implicit cross-modal alignment achieved during generative pretraining, where the language decoder learns to exploit multimodal signals within a shared representation space for generating unimodal outputs. Through analysis of anisotropy and kernel similarity structure, we empirically confirm that latent alignment emerges within MLLM representations, allowing CL to serve as a lightweight refinement stage. Leveraging this insight, we propose a Language-Centric Omnimodal Embedding framework, termed LCO-EMB. Extensive experiments across diverse backbones and benchmarks demonstrate its effectiveness, achieving state-of-the-art performance across modalities. Furthermore, we identify a Generation-Representation Scaling Law (GRSL), showing that the representational capabilities gained through contrastive refinement scales positively with the MLLM's generative capabilities. This suggests that improving generative abilities evolves as an effective paradigm for enhancing representation quality. We provide a theoretical explanation of GRSL, which formally links the MLLM's generative quality to the upper bound on its representation performance, and validate it on a challenging, low-resource visual-document retrieval task, showing that continual generative pretraining before CL can further enhance the potential of a model's embedding capabilities. 1 Recent approaches utilize autoregressive multimodal large language models (MLLMs) as the backbone models, followed by CL fine-tuning, to enhance representational capabilities, leading to improved performance on these complicated tasks [8, 39, 85] . However, the underlying reasons for the performance advantages of MLLM-based embedding approaches over traditional CLIP-based ones remain underexplored. This represents a critical research gap in understanding the limitations of CLIP-style models and the specific strengths MLLMs bring to these challenging scenarios.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4081f5b9-6560-424a-b56e-2f27be827755

Cited by top-tier papers2

Ask how each one uses it

Builds on29

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines