Zeros can be Informative: Masked Binary U-Net for Image Segmentation on Tensor Cores
Chunshu Wu, Ruibing Song, Sushant Kondguli, Tony Geng, Ang Li
Abstract
Real-time image segmentation is a key enabler for AR/VR, robotics, drones, and autonomous systems, where tight accuracy, latency, and energy budgets must be met on resource-constrained edge devices. While U-Net offers a favorable balance of accuracy and efficiency compared to large transformer-based models, achieving real-time performance on high-resolution input remains challenging due to compute, memory, and power limits. Extreme quantization, particularly binary networks, is appealing for its hardware-friendly operations. However, two obstacles limit practicality: (1) severe accuracy degradation, and (2) a lack of end-to-end implementations that deliver efficiency on general-purpose GPUs. We make two empirical observations that guide our design. (1) An explicit zero state is essential: training with zero masking to binary U-Net weights yields noticeable sparsity. (2) Quantization sensitivity is uniform across layers. Motivated by these findings, we introduce Masked Binary U-Net (MBU-Net), obtained through a cost-aware masking strategy that prioritizes masking where it yields the highest accuracy-per-cost, reconciling accuracy with near-binary efficiency. To realize these gains in practice, we develop a GPU execution framework that maps MBU-Net to Tensor Cores via a subtractive bit-encoding scheme, efficiently implementing masked binary weights with binary activations. This design leverages native binary Tensor Core BMMA instructions, enabling high throughput and energy savings on widely available GPUs. Across 3 segmentation benchmarks, MBU-Net attains near full-precision accuracy (3% average drop) while delivering 2.04× speedup and 3.54× energy reductions over a 16-bit floating point U-Net. Compared to large, over-parameterized architectures such as Vision Transformer [1], U-Net [2] is substantially cheaper in compute, memory, and energy for dense prediction, making it a more promising choice for edge scenarios. Additionally, U-Net's characteristic encoder-decoder structure with skip connections has demonstrated high effectiveness for pixellevel tasks [3, 4] , preserving spatial details and facilitating accurate pixel-level predictions. Nevertheless, running a U-Net on high-resolution video in real time can still exceed the strict power, latency, and memory limits of many edge devices -further optimizations are necessary to reconcile accuracy with deployability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- FreeU: Free Lunch in Diffusion U-NetChenyang Si, Ziqi Huang, Yuming Jiang, Ziwei LiuCVPR 2024 · 111 citations
Related papers
- Towards Real-Time Segmentation on the EdgeYanyu Li, Changdi Yang, Pu Zhao, Geng Yuan et al.AAAI 2023 · 19 citations
- BiMatting: Efficient Video Matting via BinarizationHaotong Qin, Lei Ke, Xudong Ma, Martin Danelljan et al.NeurIPS 2023 · 28 citations
- AQD: Towards Accurate Quantized Object DetectionPeng Chen, Jing Liu, Bohan Zhuang, Mingkui Tan et al.CVPR 2021
- Fully Quantized Image Super-Resolution NetworksHu Wang, Peng Chen, Bohan Zhuang, Chunhua ShenACM MM 2021 · 26 citations
- Latency-aware Spatial-wise Dynamic NetworksYizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue et al.NeurIPS 2022 · 30 citations
