EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense Prediction
Han Cai, Junyan Li, Muyan Hu, Chuang Gan, Song Han
Abstract
High-resolution dense prediction enables many appealing real-world applications, such as computational photography, autonomous driving, etc. However, the vast computational cost makes deploying state-of-the-art highresolution dense prediction models on hardware devices difficult. This work presents EfficientViT, a new family of high-resolution vision models with novel lightweight multiscale attention. Unlike prior high-resolution dense prediction models that rely on heavy self-attention, hardwareinefficient large-kernel convolution, or complicated topology structure to obtain good performances, our lightweight multi-scale attention achieves a global receptive field and multi-scale learning (two critical features for highresolution dense prediction) with only lightweight and hardware-efficient operations. As such, EfficientViT delivers remarkable performance gains over previous state-ofthe-art high-resolution dense prediction models with significant speedup on diverse hardware platforms, including mobile CPU, edge GPU, and cloud GPU. Without performance loss on Cityscapes, our EfficientViT provides up to 8.8× and 3.8× GPU latency reduction over SegFormer and SegNeXt, respectively. For super-resolution, EfficientViT provides up to 6.4× speedup over Restormer while providing 0.11dB gain in PSNR.
SegFormer Untitled 1 Untitled 2 25 80.5 50.5 79.8 Untitled 3 74 82.1 124.6 81.3 243.7 78.5 Untitled 4 179 83.0 275.7 82.6 717.1 81.0
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 875b78e4-74a6-4d80-be37-879adf05896fCited by top-tier papers49
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion ModelsLvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein et al.NeurIPS 2025 · 132 citations
- Radial Attention: 𝒪(n log n) Sparse Attention with Energy Decay for Long Video GenerationXingyang Li, Muyang Li, Tianle Cai, Haocheng Xi et al.NeurIPS 2025 · 66 citations
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchYuxian Gu, Qinghao Hu, Haocheng Xi, Junyu Chen et al.NeurIPS 2025 · 39 citations
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationFenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma et al.ACM MM 2025 · 29 citations
- MetaUAS: Universal Anomaly Segmentation with One-Prompt Meta-LearningBin-Bin GaoNeurIPS 2024 · 18 citations
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
Related papers
- SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision ApplicationsAbdelrahman M. Shaker, Muhammad Maaz, Hanoona Abdul Rasheed, Salman H. Khan et al.ICCV 2023 · 213 citations
- EfficientViT: Memory Efficient Vision Transformer with Cascaded Group AttentionXinyu Liu, Houwen Peng, Ningxin Zheng, Yuqing Yang et al.CVPR 2023
- Multi-Scale High-Resolution Vision Transformer for Semantic SegmentationJiaqi Gu, Hyoukjun Kwon, Dilin Wang, Wei Ye et al.CVPR 2022 · 236 citations
- SpeedDETR: Speed-aware Transformers for End-to-end Object DetectionPeiyan Dong, Zhenglun Kong, Xin Meng, Peng Zhang et al.ICML 2023 · 21 citations
- Efficient Modulation for Vision NetworksXu Ma, Xiyang Dai, Jianwei Yang, Bin Xiao et al.ICLR 2024 · 30 citations
