EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense Prediction
Han Cai, Junyan Li, Muyan Hu, Chuang Gan, Song Han
摘要
High-resolution dense prediction enables many appealing real-world applications, such as computational photography, autonomous driving, etc. However, the vast computational cost makes deploying state-of-the-art highresolution dense prediction models on hardware devices difficult. This work presents EfficientViT, a new family of high-resolution vision models with novel lightweight multiscale attention. Unlike prior high-resolution dense prediction models that rely on heavy self-attention, hardwareinefficient large-kernel convolution, or complicated topology structure to obtain good performances, our lightweight multi-scale attention achieves a global receptive field and multi-scale learning (two critical features for highresolution dense prediction) with only lightweight and hardware-efficient operations. As such, EfficientViT delivers remarkable performance gains over previous state-ofthe-art high-resolution dense prediction models with significant speedup on diverse hardware platforms, including mobile CPU, edge GPU, and cloud GPU. Without performance loss on Cityscapes, our EfficientViT provides up to 8.8× and 3.8× GPU latency reduction over SegFormer and SegNeXt, respectively. For super-resolution, EfficientViT provides up to 6.4× speedup over Restormer while providing 0.11dB gain in PSNR.
SegFormer Untitled 1 Untitled 2 25 80.5 50.5 79.8 Untitled 3 74 82.1 124.6 81.3 243.7 78.5 Untitled 4 179 83.0 275.7 82.6 717.1 81.0
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion ModelsLvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein 等NeurIPS 2025 · 被引用 132 次
- Radial Attention: 𝒪(n log n) Sparse Attention with Energy Decay for Long Video GenerationXingyang Li, Muyang Li, Tianle Cai, Haocheng Xi 等NeurIPS 2025 · 被引用 66 次
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchYuxian Gu, Qinghao Hu, Haocheng Xi, Junyu Chen 等NeurIPS 2025 · 被引用 39 次
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationFenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma 等ACM MM 2025 · 被引用 29 次
- MetaUAS: Universal Anomaly Segmentation with One-Prompt Meta-LearningBin-Bin GaoNeurIPS 2024 · 被引用 18 次
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
相关 Paper
- SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision ApplicationsAbdelrahman M. Shaker, Muhammad Maaz, Hanoona Abdul Rasheed, Salman H. Khan 等ICCV 2023 · 被引用 213 次
- EfficientViT: Memory Efficient Vision Transformer with Cascaded Group AttentionXinyu Liu, Houwen Peng, Ningxin Zheng, Yuqing Yang 等CVPR 2023
- Multi-Scale High-Resolution Vision Transformer for Semantic SegmentationJiaqi Gu, Hyoukjun Kwon, Dilin Wang, Wei Ye 等CVPR 2022 · 被引用 236 次
- SpeedDETR: Speed-aware Transformers for End-to-end Object DetectionPeiyan Dong, Zhenglun Kong, Xin Meng, Peng Zhang 等ICML 2023 · 被引用 21 次
- Efficient Modulation for Vision NetworksXu Ma, Xiyang Dai, Jianwei Yang, Bin Xiao 等ICLR 2024 · 被引用 30 次
