NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
Loïck Chambon, Paul Couairon, Éloi Zablocki, Alexandre Boulch, Nicolas Thome, Matthieu Cord
Abstract
Vision Foundation Models (VFMs) extract spatially downsampled representations, posing challenges for pixel-level tasks. Existing upsampling approaches face a fundamental trade-off: classical filters are fast and broadly applicable but rely on fixed forms, while modern upsamplers achieve superior accuracy through learnable, VFM-specific forms at the cost of retraining for each VFM. We introduce Neighborhood Attention Filtering (NAF), which bridges this gap by learning adaptive spatial-and-content weights through Cross-Scale Neighborhood Attention and Rotary Position Embeddings (RoPE), guided solely by the high-resolution input image. NAF operates zero-shot: it upsamples features from any VFM without retraining, making it the first VFM-agnostic architecture to outperform VFM-specific upsamplers and achieve state-of-the-art performance across multiple downstream tasks. It maintains high efficiency, scaling to 2K feature maps and reconstructing intermediate-resolution maps at 18 FPS. Beyond feature upsampling, NAF demonstrates strong performance on image restoration, highlighting its versatility. Code and checkpoints are available at https://github.com/valeoai/NAF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d24880f4-9641-44ee-9bdc-6659972c5aadCited by top-tier papers1
Ask how each one uses itBuilds on17
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- CARAFE: Content-Aware ReAssembly of FEaturesJiaqi Wang, Kai Chen, Rui Xu, Ziwei Liu et al.ICCV 2019 · 842 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Learning to Upsample by Learning to SampleWenze Liu, Hao Lu, Hongtao Fu, Zhiguo CaoICCV 2023 · 518 citations
Related papers
- LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation ModelsHaiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger et al.ICCV 2025 · 7 citations
- JAFAR: Jack up Any Feature at Any ResolutionPaul Couairon, Loïck Chambon, Louis Serrano, Jean-Emmanuel Haugeard et al.NeurIPS 2025 · 26 citations
- AnyUp: Universal Feature UpsamplingThomas Wimmer, Prune Truong, Marie-Julie Rakotosaona, Michael Oechsle et al.ICLR 2026 · 29 citations
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local AttendersMatthew Walmer, Saksham Suri, Anirud Aggarwal, Abhinav ShrivastavaCVPR 2026 · 2 citations
- Compass-RoPE: Isotropic Rotary Position Embeddings for Vision TransformersChengxi Min, Wei Wang, Yao ZhaoICML 2026
