JAFAR: Jack up Any Feature at Any Resolution
Paul Couairon, Loïck Chambon, Louis Serrano, Jean-Emmanuel Haugeard, Matthieu Cord, Nicolas Thome
摘要
Foundation Vision Encoders have become essential for a wide range of dense vision tasks. However, their low-resolution spatial feature outputs necessitate feature upsampling to produce the high-resolution modalities required for downstream tasks. In this work, we introduce JAFAR, a lightweight and flexible feature upsampler that enhances the spatial resolution of visual features from any Foundation Vision Encoder to an arbitrary target resolution. JAFAR employs an attention-based module designed to promote semantic alignment between high-resolution queries, derived from low-level image features, and semantically enriched low-resolution keys, using Spatial Feature Transform (SFT) modulation. Notably, despite the absence of high-resolution supervision, we demonstrate that learning at low upsampling ratios and resolutions generalizes remarkably well to significantly higher output scales. Extensive experiments show that JAFAR effectively recovers fine-grained spatial details and consistently outperforms existing feature upsampling methods across a diverse set of downstream tasks. Project page at https://jafar-upsampler.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AnyUp: Universal Feature UpsamplingThomas Wimmer, Prune Truong, Marie-Julie Rakotosaona, Michael Oechsle 等ICLR 2026 · 被引用 29 次
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature UpsamplingMinseok Seo, Mark Hamilton, Changick KimCVPR 2026 · 被引用 9 次
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention FilteringLoïck Chambon, Paul Couairon, Éloi Zablocki, Alexandre Boulch 等CVPR 2026 · 被引用 6 次
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local AttendersMatthew Walmer, Saksham Suri, Anirud Aggarwal, Abhinav ShrivastavaCVPR 2026 · 被引用 2 次
- IdEst: Assessing Self-Supervised Learning Representations via Intrinsic DimensionJulie Mordacq, Vicky Kalogeiton, Steve OudotICML 2026 · 被引用 1 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
相关 Paper
- LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation ModelsHaiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger 等ICCV 2025 · 被引用 7 次
- SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary SegmentationNaomi Kombol, Ivan Martinovic, Sinisa Segvic, Giorgos ToliasCVPR 2026
- FeatSharp: Your Vision Model Features, SharperMike Ranzinger, Greg Heinrich, Pavlo Molchanov, Bryan Catanzaro 等ICML 2025
- Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-ResolutionYixuan Sun, Dongyang Zhao, Zhangyue Yin, Yiwen Huang 等CVPR 2023
- FeatUp: A Model-Agnostic Framework for Features at Any ResolutionStephanie Fu, Mark Hamilton, Laura E. Brandt, Axel Feldmann 等ICLR 2024 · 被引用 117 次
