The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-Modal Divergence
Yichao Cai, Zhen Zhang, Yuhang Liu, Javen Qinfeng Shi
Abstract
While InfoNCE underlies modern contrastive learning, its geometric mechanisms remain under-characterized beyond the canonical alignment-uniformity decomposition. We develop a measure-theoretic framework in which representation measures evolve on a fixed embedding manifold. In the large-batch limit, we prove value and gradient consistency, linking the stochastic objective to explicit deterministic energy landscapes and revealing a geometric bifurcation between unimodal and symmetric multimodal regimes. In the unimodal case, the intrinsic energy is strictly convex and admits a unique Gibbs equilibrium, showing that entropy acts as a tie-breaker within the aligned basin. In the multimodal case, the intrinsic geometry becomes cross-coupled and contains a persistent negative symmetric divergence term: each modality's marginal reshapes the effective landscape of the other, allowing strong pairwise alignment to coexist with a persistent modality gap. Controlled synthetic experiments and analyses of pretrained CLIP representations support these predictions. Overall, our results shift the analytical lens from pointwise discrimination to population geometry, showing that pairwise alignment alone is insufficient to control cross-modal marginal structure. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6380c60b-2c17-48e2-9646-369fb7c2eef7Cited by top-tier papers1
Ask how each one uses itBuilds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- Towards Uniformity and Alignment for Multimodal Representation LearningWenzhe Yin, Pan Zhou, Zehao Xiao, Jie Liu et al.ICML 2026 · 4 citations
- The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation LearningJustinas Zaliaduonis, Patrick Putzky, Till Richter, Sergios GatidisICML 2026
- The Double-Ellipsoid Geometry of CLIPMeir Yossef Levi, Guy GilboaICML 2025
- Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity MetricToshimitsu Uesaka, Taiji Suzuki, Yuhta Takida, Chieh-Hsin Lai et al.ICLR 2025
- Understanding and Constructing Latent Modality Structures in Multi-Modal Representation LearningQian Jiang, Changyou Chen, Han Zhao, Liqun Chen et al.CVPR 2023
