Adaptive Geometry Routing for Vision-Language Understanding
Sarthak Srivastava, Kathy Wu
Abstract
Vision language models face a fundamental geometry trade-off: Euclidean representations excel at instance-level discrimination, while hyperbolic representations naturally encode semantic hierarchies. Hybrid training is challenging because one geometry may dominate early, leaving the other under-trained failure mode we term geometry dominance. We introduce Adaptive Geometry Routing (AGR), a framework that addresses this via a novel four-phase training curriculum : (1) Isolation hyperbolic-only training stabilizes hierarchical structure; (2) Shadow router learns mixing patterns using only hyperbolic signals; (3) Soft Launch Euclidean scores gradually become visible; (4) Adaptive full dual-geometry routing. This phased coordination of router activation (?) and Euclidean visibility (?) prevents early dominance while enabling data-driven geometry selection. Built on a shared backbone with lightweight LoRA-adapted heads and bounded residual corrections, AGR discovers that hyperbolic geometry is preferred by default (85% weight), with routing adapting semantically abstract queries route more hyperbolic, attribute-rich queries shift toward Euclidean. On ViT-B, AGR achieves 38.8% COCO T2I R@5 (+6.5pp over MERU), 64.7% Flickr30K I2T R@5 (+11.3pp), and 32.6% ImageNet accuracy (+9.3pp), demonstrating that phased curriculum training enables stable hybrid geometry learning for vision-language understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d120b997-d2c2-4c6a-8a78-b383e1a33010Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model AdaptationXi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu et al.ACL 2026 · 6 citations
- Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent SpaceCheng Yan, Wuyang Zhang, Zhiyuan Ning, Fan Xu et al.AAAI 2026
- VladVA: Discriminative Fine-tuning of LVLMsYassine Ouali, Adrian Bulat, Alexandros Xenos, Anestis Zaganidis et al.CVPR 2025
- Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context LearningNiloufar Alipour Talemi, Hossein Kashiani, Fatemeh AfghahICML 2026
- CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language DetectionZhipeng Liu, Chunbo LuoCVPR 2026 · 1 citation
