HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
Zelin Peng, Zhengqin Xu, Qingyang Liu, Xiaokang Yang, Wei Shen
Abstract
Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computational resources (e.g., thousands of GPUs) for training to achieve cross-modal alignment at multi-granularity levels. We argue that a key source of this inefficiency lies in the vision encoders they widely equip with, e.g., CLIP and SAM, which lack the alignment with language at multi-granularity levels. To address this issue, in this paper, we leverage hyperbolic space, which inherently models hierarchical levels and thus provides a principled framework for bridging the granularity gap between visual and textual modalities at an arbitrary granularity level. Concretely, we propose an efficient training paradigm for MLLMs, dubbed as HyperET, which can optimize visual representations to align with their textual counterparts at an arbitrary granularity level through dynamic hyperbolic radius adjustment in hyperbolic space. HyperET employs learnable matrices with Möbius multiplication operations, implemented via three effective configurations: diagonal scaling matrices, block-diagonal matrices, and banded matrices, providing a flexible yet efficient parametrization strategy. Comprehensive experiments across multiple MLLM benchmarks demonstrate that HyperET consistently improves both existing pre-training and fine-tuning MLLMs clearly with less than 1% additional parameters. Code is available at https://github.com/godlin-sjtu/HyperET
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9fb424e-42c3-48a0-81ff-9d543574b151Cited by top-tier papers6
- Rejection Mixing: Fast Semantic Propagation of Mask Tokens for Efficient DLLM InferenceYushi Ye, Feng Hong, Huangjie Zheng, Xu Chen et al.CVPR 2026 · 5 citations
- Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion TransformersXinyu Peng, Han Li, Yuyang Huang, Ziyang Zheng et al.CVPR 2026 · 4 citations
- Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual ReasonersQingyang Liu, Bingjie Gao, Canmiao Fu, Zhipeng Huang et al.ICML 2026 · 1 citation
- Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic SegmentationXianglin Qiu, Jian Wang, Xiaolei Wang, Zhen Zhang et al.CVPR 2026
- Hyperbolic Gramian Volumes for Multimodal AlignmentSaiyang Na, Feng Jiang, Qifeng Zhou, Wenliang Zhong et al.CVPR 2026
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic SpaceZelin Peng, Zhengqin Xu, Zhilin Zeng, Changsong Wen et al.CVPR 2025
- ROSE: Rotate Your Large Language Model to SeeTongtian Yue, Xuange Gao, Longteng Guo, Zijia Zhao et al.CVPR 2026
- Hyperbolic Image-text RepresentationsKaran Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson et al.ICML 2023 · 137 citations
- Visual Perception by Large Language Model's WeightsFeipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang et al.NeurIPS 2024 · 24 citations
- Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal RepresentationsJeonghyeon Kim, Sangheum HwangCVPR 2025
