Zeros can be Informative: Masked Binary U-Net for Image Segmentation on Tensor Cores
Chunshu Wu, Ruibing Song, Sushant Kondguli, Tony Geng, Ang Li
摘要
Real-time image segmentation is a key enabler for AR/VR, robotics, drones, and autonomous systems, where tight accuracy, latency, and energy budgets must be met on resource-constrained edge devices. While U-Net offers a favorable balance of accuracy and efficiency compared to large transformer-based models, achieving real-time performance on high-resolution input remains challenging due to compute, memory, and power limits. Extreme quantization, particularly binary networks, is appealing for its hardware-friendly operations. However, two obstacles limit practicality: (1) severe accuracy degradation, and (2) a lack of end-to-end implementations that deliver efficiency on general-purpose GPUs. We make two empirical observations that guide our design. (1) An explicit zero state is essential: training with zero masking to binary U-Net weights yields noticeable sparsity. (2) Quantization sensitivity is uniform across layers. Motivated by these findings, we introduce Masked Binary U-Net (MBU-Net), obtained through a cost-aware masking strategy that prioritizes masking where it yields the highest accuracy-per-cost, reconciling accuracy with near-binary efficiency. To realize these gains in practice, we develop a GPU execution framework that maps MBU-Net to Tensor Cores via a subtractive bit-encoding scheme, efficiently implementing masked binary weights with binary activations. This design leverages native binary Tensor Core BMMA instructions, enabling high throughput and energy savings on widely available GPUs. Across 3 segmentation benchmarks, MBU-Net attains near full-precision accuracy (3% average drop) while delivering 2.04× speedup and 3.54× energy reductions over a 16-bit floating point U-Net. Compared to large, over-parameterized architectures such as Vision Transformer [1], U-Net [2] is substantially cheaper in compute, memory, and energy for dense prediction, making it a more promising choice for edge scenarios. Additionally, U-Net's characteristic encoder-decoder structure with skip connections has demonstrated high effectiveness for pixellevel tasks [3, 4] , preserving spatial details and facilitating accurate pixel-level predictions. Nevertheless, running a U-Net on high-resolution video in real time can still exceed the strict power, latency, and memory limits of many edge devices -further optimizations are necessary to reconcile accuracy with deployability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- FreeU: Free Lunch in Diffusion U-NetChenyang Si, Ziqi Huang, Yuming Jiang, Ziwei LiuCVPR 2024 · 被引用 111 次
相关 Paper
- Towards Real-Time Segmentation on the EdgeYanyu Li, Changdi Yang, Pu Zhao, Geng Yuan 等AAAI 2023 · 被引用 19 次
- BiMatting: Efficient Video Matting via BinarizationHaotong Qin, Lei Ke, Xudong Ma, Martin Danelljan 等NeurIPS 2023 · 被引用 28 次
- AQD: Towards Accurate Quantized Object DetectionPeng Chen, Jing Liu, Bohan Zhuang, Mingkui Tan 等CVPR 2021
- Fully Quantized Image Super-Resolution NetworksHu Wang, Peng Chen, Bohan Zhuang, Chunhua ShenACM MM 2021 · 被引用 26 次
- Latency-aware Spatial-wise Dynamic NetworksYizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue 等NeurIPS 2022 · 被引用 30 次
