FeatSharp: Your Vision Model Features, Sharper
Mike Ranzinger, Greg Heinrich, Pavlo Molchanov, Bryan Catanzaro, Andrew Tao
摘要
The feature maps of vision encoders are fundamental to myriad modern AI tasks, ranging from core perception algorithms (e.g. semantic segmentation, object detection, depth perception, etc.) to modern multimodal understanding in visionlanguage models (VLMs). Currently, in computer vision, the frontier of general purpose vision backbones is Vision Transformers (ViT), typically trained using contrastive loss (e.g. CLIP). A key problem with most off-the-shelf ViTs, particularly CLIP, is that these models are inflexibly low resolution. Most run at 224 × 224px, while the "high-resolution" versions are around 378 -448px, but still inflexible. We introduce a novel method to coherently and cheaply upsample the feature maps of low-resolution vision encoders while picking up on fine-grained details that would otherwise be lost due to resolution. We demonstrate the effectiveness of this approach on core perception tasks as well as within agglomerative model training using RADIO as a way of providing richer targets for distillation. Code available at https://github.com/NVlabs/FeatSharp .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AnyUp: Universal Feature UpsamplingThomas Wimmer, Prune Truong, Marie-Julie Rakotosaona, Michael Oechsle 等ICLR 2026 · 被引用 29 次
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic EncodingYueming Xu, Jiahui Zhang, Ze Huang, Yurui Chen 等ICLR 2026 · 被引用 8 次
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention FilteringLoïck Chambon, Paul Couairon, Éloi Zablocki, Alexandre Boulch 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation ModelsHaiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger 等ICCV 2025 · 被引用 7 次
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense PredictionJuan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin 等ICCV 2025
- SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary SegmentationNaomi Kombol, Ivan Martinovic, Sinisa Segvic, Giorgos ToliasCVPR 2026
- MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any ResolutionWenzhuo Liu, Fei Zhu, Shijie Ma, Cheng-Lin LiuNeurIPS 2024 · 被引用 17 次
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 等ICLR 2024 · 被引用 129 次
