TroL: Traversal of Layers for Large Language and Vision Models
Byung-Kwan Lee, Sangyun Chung, Chae Won Kim, Beomchan Park, Yong Man Ro
摘要
Large language and vision models (LLVMs) have been driven by the generalization power of large language models (LLMs) and the advent of visual instruction tuning.Along with scaling them up directly, these models enable LLVMs to showcase powerful vision language (VL) performances by covering diverse tasks via natural language instructions.However, existing open-source LLVMs that perform comparably to closed-source LLVMs such as GPT-4V are often considered too large (e.g., 26B, 34B, and 110B parameters), having a larger number of layers.These large models demand costly, high-end resources for both training and inference.To address this issue, we present a new efficient LLVM family with 1.8B, 3.8B, and 7B LLM model sizes, Traversal of Layers ( TroL), which enables the reuse of layers in a token-wise manner.This layer traversing technique simulates the effect of looking back and retracing the answering stream while increasing the number of forward propagation layers without physically adding more layers.We demonstrate that TroL employs a simple layer traversing approach yet efficiently outperforms the open-source LLVMs with larger model sizes and rivals the performances of the closed-source LLVMs with substantial sizes.Code is available in https://github.com/ByungKwanLee/TroL.MM1-MoE Monkey-Qwen LLaVA-Next-LLaMA3 LLaVA-NeXT-Mistral MiniGemini-HD-Vicuna InternVL1.5-InternLM2-ChatInternVL1.5-InternLM2-Chat
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon 等ICLR 2026 · 被引用 13 次
- Unified Reinforcement and Imitation Learning for Vision-Language ModelsByung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang 等NeurIPS 2025 · 被引用 12 次
- Masking Teacher and Reinforcing Student for Distilling Vision-Language ModelsByung-Kwan Lee, Yu-Chiang Frank Wang, Ryo HachiumaCVPR 2026 · 被引用 7 次
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language ModelsYoung-Jun Lee, Byung-Kwan Lee, Jianshu Zhang, Yechan Hwang 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Meteor: Mamba-based Traversal of Rationale for Large Language and Vision ModelsByung-Kwan Lee, Chae Won Kim, Beomchan Park, Yong Man RoNeurIPS 2024 · 被引用 37 次
- Accelerating Multimodal Large Language Models by Searching Optimal Vision Token ReductionShiyu Zhao, Zhenting Wang, Felix Juefei-Xu, Xide Xia 等CVPR 2025
- VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision ComputationShiwei Wu, Joya Chen, Kevin Qinghong Lin, Qimeng Wang 等NeurIPS 2024 · 被引用 78 次
- Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and InferenceSiyuan Wang, Dianyi Wang, Chengxing Zhou, Zejun Li 等ACL 2025 · 被引用 3 次
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao 等ACL 2025
