Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Max Ehrlich, Tong Lu, Limin Wang
摘要
We introduce Eagle 2.5, a family of frontier vision-language models (VLMs) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and high-resolution image understanding, introducing a generalist framework for both tasks. The proposed training framework incorporates Automatic Degrade Sampling and Image Area Preservation, two techniques that preserve contextual integrity and visual details. The framework also includes numerous efficiency optimizations in the pipeline for long-context data training. Finally, we propose Eagle-Video-110K, a novel dataset that integrates both story-level and clip-level annotations, facilitating long-video understanding. Eagle 2.5 demonstrates substantial improvements on long-context multimodal benchmarks, providing a robust solution to the limitations of existing VLMs. Notably, our best model Eagle 2.5-8B achieves 72.4% on Video-MME with 512 input frames, matching the results of top-tier commercial model such as GPT-4o and large-scale open-source models like Qwen2.5-VL-72B and InternVL2.5-78B.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningChi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang 等NeurIPS 2025 · 被引用 179 次
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLMHanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang 等ICLR 2026 · 被引用 64 次
- VideoITG: Multimodal Video Understanding with Instructed Temporal GroundingShihao Wang, Guo Chen, De-An Huang, Zhiqi Li 等CVPR 2026 · 被引用 35 次
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought ReasoningQing Jiang, Xingyu Chen, Zhaoyang Zeng, Junzhi Yu 等ICLR 2026 · 被引用 25 次
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng 等ICLR 2026 · 被引用 24 次
它引用的顶会 Paper52
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- EAGLE: Egocentric AGgregated Language-video EngineJing Bi, Yunlong Tang, Luchuan Song, Ali Vosoughi 等ACM MM 2024 · 被引用 3 次
- ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction TuningRui Wang, Bohao Li, Xiyang Dai, Jianwei Yang 等EMNLP 2025
- V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position EncodingJunqi Ge, Ziyi Chen, Jintao Lin, Jinguo Zhu 等ICCV 2025 · 被引用 5 次
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li 等CVPR 2025
- Visual Context Window Extension: A New Perspective for Long Video UnderstandingHongchen Wei, Zhenzhong ChenACM MM 2025
