Lune

CVPR2025Top-tier venue

NVILA: Efficient Frontier Visual Language Models

Zhijian Liu, Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Haotian Tang

2025Year
72Top-tier citations

Abstract

Decoding Speed (token/s) (1.9-5.1X Faster) (Pre-fill: 1.6-2.2X Faster / Decode: 1.2-2.8X Faster) 5.1X Faster 2.2X Faster 2.8X Faster AI2D C h a rt Q A D oc VQ A In fo VQ A M a th V is ta MMMU R e a lW o rl d Q A S EE D (im ag e) Te xt VQ A V Q A v 2 (c) Accuracy on image and video benchmarks (On-par or superior accuracy on all benchmarks)

Efficient Frontier VLMs. (a) NVILA trains image and video models 5.1→ and 1.9→ faster, respectively, than LLaVA-OneVision (OV), which is the only baseline model with publicly disclosed training costs. (b) Against Qwen2-VL, NVILA achieves a 1.6-2.2→ measured speedup in the pre-filling stage and a 1.2-2.8→ speedup during the decoding stage. (c) NVILA's efficiency is achieved without compromising accuracy; in fact, it delivers comparable or even superior accuracy across image and video benchmarks. All models in this table have 8B parameters. Training time in (a) is measured using NVIDIA H100 GPUs, while inference speed in (b) is measured using a single NVIDIA GeForce RTX 4090 GPU. Accuracy numbers in (c) are normalized relative to the highest score for each benchmark.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 585a14a6-3aef-499a-87bf-2c6a41e7f651

Cited by top-tier papers72

Ask how each one uses it

Builds on36

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines