Lune

ICLR2026Top-tier venue

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

Sharut Gupta, Shobhita Sundaram, Chenyu Wang, Stefanie Jegelka, Phillip Isola

2026Year
3Top-tier citations

Abstract

Traditional multimodal learners find unified representations for tasks like visual question answering, but rely heavily on large paired datasets. However, an overlooked yet potentially powerful question is: can one leverage auxiliary unpaired\textit{unpaired} multimodal data to directly enhance representation learning in a target\textit{target} modality? We introduce UML\textbf{UML}: U\textbf{U}npaired M\textbf{M}ultimodal L\textbf{L}earner, a modality-agnostic training paradigm in which a single model alternately processes inputs from different modalities while sharing parameters across them. This design exploits the assumption that different modalities are projections of a shared underlying reality, allowing the model to benefit from cross-modal structure without requiring explicit pairs. Theoretically, under linear data-generating assumptions, we show that unpaired auxiliary data can yield representations strictly more informative about the world than unimodal training. Empirically, we show that incorporating unpaired data that share underlying semantic information from auxiliary modalities—such as text, audio, or images—consistently improves downstream performance across diverse unimodal targets such as image and audio. Our project page: https://unpaired-multimodal.github.io/

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 109896b7-35d5-416b-99ad-d06f0762f87b

Cited by top-tier papers3

Ask how each one uses it

Builds on42

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines