Soft Superpixel Neighborhood Attention
Kent W. Gauen, Stanley H. Chan
Abstract
Images contain objects with deformable boundaries, such as the contours of a human face, yet attention operators act on square windows. This mixes features from perceptually unrelated regions, which can degrade the quality of a denoiser. One can exclude pixels using an estimate of perceptual groupings, such as superpixels, but the naive use of superpixels can be theoretically and empirically worse than standard attention. Using superpixel probabilities rather than superpixel assignments, this paper proposes soft superpixel neighborhood attention (SNA) which interpolates between the existing neighborhood attention and the naive superpixel neighborhood attention. This paper presents theoretical results showing SNA is the optimal denoiser under a latent superpixel model. SNA outperforms alternative lo-cal attention modules on image denoising, and we compare the superpixels learned from denoising with those learned with superpixel supervision. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e1e4f91-2d99-4b84-8d45-d0539894cbb6Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- XCiT: Cross-Covariance Image TransformersAlaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski et al.NeurIPS 2021 · 692 citations
Related papers
- FSNet: Frequency Domain Guided Superpixel Segmentation Network for Complex ScenesHua Li, Junyan Liang, Wenjie Li, Wenhui WuACM MM 2023 · 9 citations
- Robust Superpixel-Guided Attentional Adversarial AttackXiaoyi Dong, Jiangfan Han, Dongdong Chen, Jiayang Liu et al.CVPR 2020
- Lightweight Image Super-Resolution with Superpixel Token InteractionAiping Zhang, Wenqi Ren, Yi Liu, Xiaochun CaoICCV 2023 · 59 citations
- Spatially Adaptive Self-Supervised Learning for Real-World Image DenoisingJunyi Li, Zhilu Zhang, Xiaoyu Liu, Chaoyu Feng et al.CVPR 2023
- Interpreting Super-Resolution Networks With Local Attribution MapsJinjin Gu, Chao DongCVPR 2021
