LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models
Haiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger, Dan Zhang
Abstract
Vision foundation models (VFMs) such as DINOv2 and CLIP have achieved impressive results on various downstream tasks, but their limited feature resolution hampers performance in applications requiring pixel-level understanding. Feature upsampling offers a promising direction to address this challenge. In this work, we identify two critical factors for enhancing feature upsampling: the upsampler architecture and the training objective. For the upsampler architecture, we introduce a coordinate-based crossattention transformer that integrates the high-resolution images with coordinates and low-resolution VFM features to generate sharp, high-quality features. For the training objective, we propose constructing high-resolution pseudogroundtruth features by leveraging class-agnostic masks and self-distillation. Our approach effectively captures fine-grained details and adapts flexibly to various input and feature resolutions. Through experiments, we demonstrate that our approach significantly outperforms existing feature upsampling techniques across various downstream tasks. Our code is released at https://github.com/ andrehuang/loftup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f1f15ce-fb46-4fa9-9df6-77ccfda0e0e7Cited by top-tier papers13
- AnyUp: Universal Feature UpsamplingThomas Wimmer, Prune Truong, Marie-Julie Rakotosaona, Michael Oechsle et al.ICLR 2026 · 29 citations
- JAFAR: Jack up Any Feature at Any ResolutionPaul Couairon, Loïck Chambon, Louis Serrano, Jean-Emmanuel Haugeard et al.NeurIPS 2025 · 26 citations
- CoWTracker: Tracking by Warping instead of CorrelationZihang Lai, Eldar Insafutdinov, Edgar Sucar, Andrea VedaldiCVPR 2026 · 12 citations
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature UpsamplingMinseok Seo, Mark Hamilton, Changick KimCVPR 2026 · 9 citations
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention FilteringLoïck Chambon, Paul Couairon, Éloi Zablocki, Alexandre Boulch et al.CVPR 2026 · 6 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- CARAFE: Content-Aware ReAssembly of FEaturesJiaqi Wang, Kai Chen, Rui Xu, Ziwei Liu et al.ICCV 2019 · 842 citations
Related papers
- FeatSharp: Your Vision Model Features, SharperMike Ranzinger, Greg Heinrich, Pavlo Molchanov, Bryan Catanzaro et al.ICML 2025
- Improving CLIP Fine-tuning PerformanceYixuan Wei, Han Hu, Zhenda Xie, Ze Liu et al.ICCV 2023 · 24 citations
- FlowFeat: Pixel-Dense Embedding of Motion ProfilesNikita Araslanov, Anna Sonnweber, Daniel CremersNeurIPS 2025 · 3 citations
- RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation ModelsGreg Heinrich, Mike Ranzinger, Hongxu Yin, Yao Lu et al.CVPR 2025
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense PredictionJuan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin et al.ICCV 2025
