Global and Local Entailment Learning for Natural World Imagery
Srikumar Sastry, Aayush Dhakal, Eric Xing, Subash Khanal, Nathan Jacobs
摘要
Learning the hierarchical structure of data in vision-language models is a significant challenge. Previous works have attempted to address this challenge by employing entailment learning. However, these approaches fail to model the transitive nature of entailment explicitly, which establishes the relationship between order and semantics within a representation space. In this work, we introduce Radial Cross-Modal Embeddings (RCME), a framework that enables the explicit modeling of transitivity-enforced entailment. Our proposed framework optimizes for the partial order of concepts within vision-language models. By leveraging our framework, we develop a hierarchical vision-language foundation model capable of representing the hierarchy in the Tree of Life. Our experiments on hierarchical species classification and hierarchical retrieval tasks demonstrate the enhanced performance of our models compared to the existing state-of-the-art models. Our code and models are open-sourced at https://vishu26.github.io/RCME/index.html.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Modality Alignment across Trees on Heterogeneous Hyperbolic ManifoldsWei Wu, Xiaomeng Fan, Yuwei Wu, Zhi Gao 等ICLR 2026 · 被引用 3 次
- Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept LearningSahil Mishra, Srinitish Srinivasan, Sourish Dasgupta, Tanmoy ChakrabortyICML 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 被引用 303 次
相关 Paper
- Compositional Entailment Learning for Hyperbolic Vision-Language ModelsAvik Pal, Max van Spengler, Guido Maria D'Amely di Melendugno, Alessandro Flaborea 等ICLR 2025
- Learning Visual Hierarchies in Hyperbolic Space for Image RetrievalZiwei Wang, Sameera Ramasinghe, Chenchen Hu, Julien Monteil 等ICCV 2025 · 被引用 4 次
- VL-KGE: Vision-Language Models Meet Knowledge Graph EmbeddingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg 等WWW 2026 · 被引用 2 次
- PHyCLIP: -Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation LearningDaiki Yoshikawa, Takashi MatsubaraICLR 2026
- Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal ModelsHulingxiao He, Zhi Tan, Yuxin PengICML 2026
