Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
Minseok Seo, Mark Hamilton, Changick Kim
Abstract
We present Upsample Anything, a lightweight test-time optimization (TTO) framework that restores low-resolution features to high-resolution, pixel-wise outputs without any training. Although Vision Foundation Models demonstrate strong generalization across diverse downstream tasks, their representations are typically downsampled by 14x/16x (e.g., ViT), which limits their direct use in pixel-level applications. Existing feature upsampling approaches depend on dataset-specific retraining or heavy implicit optimization, restricting scalability and generalization. Upsample Anything addresses these issues through a simple per-image optimization that learns an anisotropic Gaussian kernel combining spatial and range cues, effectively bridging Gaussian Splatting and Joint Bilateral Upsampling. The learned kernel acts as a universal, edge-aware operator that transfers seamlessly across architectures and modalities, enabling precise high-resolution reconstruction of features, depth, or probability maps. It runs in only per 224x224 image and achieves state-of-the-art performance on semantic segmentation, depth estimation, and both depth and probability map upsampling. Project page: https://seominseok0429.github.io/Upsample-Anything/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bc6fe1c-3017-4aeb-9e5f-7f4508bd8af0Cited by top-tier papers1
Ask how each one uses itBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- AnyUp: Universal Feature UpsamplingThomas Wimmer, Prune Truong, Marie-Julie Rakotosaona, Michael Oechsle et al.ICLR 2026 · 29 citations
- FeatUp: A Model-Agnostic Framework for Features at Any ResolutionStephanie Fu, Mark Hamilton, Laura E. Brandt, Axel Feldmann et al.ICLR 2024 · 117 citations
- JAFAR: Jack up Any Feature at Any ResolutionPaul Couairon, Loïck Chambon, Louis Serrano, Jean-Emmanuel Haugeard et al.NeurIPS 2025 · 26 citations
- UnSAMFlow: Unsupervised Optical Flow Guided by Segment Anything ModelShuai Yuan, Lei Luo, Zhuo Hui, Can Pu et al.CVPR 2024 · 7 citations
- Pixel to Gaussian: Ultra-Fast Continuous Super-Resolution with 2D Gaussian ModelingLong Peng, Anran Wu, Wenbo Li, Peizhe Xia et al.ICLR 2026 · 57 citations
