Compositional Entailment Learning for Hyperbolic Vision-Language Models
Avik Pal, Max van Spengler, Guido Maria D'Amely di Melendugno, Alessandro Flaborea, Fabio Galasso, Pascal Mettes
摘要
Image-text representation learning forms a cornerstone in vision-language models, where pairs of images and textual descriptions are contrastively aligned in a shared embedding space. Since visual and textual concepts are naturally hierarchical, recent work has shown that hyperbolic space can serve as a high-potential manifold to learn vision-language representation with strong downstream performance. In this work, for the first time we show how to fully leverage the innate hierarchical nature of hyperbolic embeddings by looking beyond individual image-text pairs. We propose Compositional Entailment Learning for hyperbolic vision-language models. The idea is that an image is not only described by a sentence but is itself a composition of multiple object boxes, each with their own textual description. Such information can be obtained freely by extracting nouns from sentences and using openly available localized grounding models. We show how to hierarchically organize images, image boxes, and their textual descriptions through contrastive and entailment-based objectives. Empirical evaluation on a hyperbolic visionlanguage model trained with millions of image-text pairs shows that the proposed compositional learning approach outperforms conventional Euclidean CLIP learning, as well as recent hyperbolic alternatives, with better zero-shot and retrieval generalization and clearly stronger hierarchical performance. Code to be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- Hyperbolic Fine-Tuning for Large Language ModelsMenglin Yang, Ram Samarth B. B., Aosong Feng, Bo Xiong 等NeurIPS 2025 · 被引用 31 次
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature ExpertsNeil He, Rishabh Anand, Hiren Madhu, Ali Maatouk 等NeurIPS 2025 · 被引用 27 次
- Hyperbolic Dataset DistillationWenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa 等NeurIPS 2025 · 被引用 17 次
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual RecognitionYuwen Tan, Yuan Qing, Boqing GongCVPR 2026 · 被引用 6 次
- GeoWorld: Geometric World ModelsZeyu Zhang, Danning Li, Ian Reid, Richard HartleyCVPR 2026 · 被引用 6 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language ModelsHayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young ChunCVPR 2026 · 被引用 2 次
- Hyperbolic Image-text RepresentationsKaran Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson 等ICML 2023 · 被引用 137 次
- PHyCLIP: -Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation LearningDaiki Yoshikawa, Takashi MatsubaraICLR 2026
- Hyperbolic Contrastive Learning for Visual Representations beyond ObjectsSongwei Ge, Shlok Mishra, Simon Kornblith, Chun-Liang Li 等CVPR 2023
- Learning Visual Hierarchies in Hyperbolic Space for Image RetrievalZiwei Wang, Sameera Ramasinghe, Chenchen Hu, Julien Monteil 等ICCV 2025 · 被引用 4 次
