Lune

NeurIPS2025顶会

Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2

Joel Valdivia Ortega, Lorenz Lamm, Franziska Eckardt, Benedikt Schworm, Marion Jasnin, Tingying Peng

2025年份
1被引次数

摘要

Vision Transformers (ViTs), such as DINOv2, achieve strong performance across domains but often repurpose low-informative patch tokens in ways that reduce the interpretability of attention and feature maps. This challenge is especially evident in medical imaging, where domain shifts can degrade both performance and transparency. In this paper, we introduce Randomized-MLP (RMLP) regularization, a contrastive learning-based method that encourages more semantically aligned representations. We use RMLPs when fine-tuning DINOv2 to both medical and natural image modalities, showing that it improves or maintains downstream performance while producing more interpretable attention maps. We also provide a mathematical analysis of RMLPs, offering insights into its role in enhancing ViT-based models and advancing our understanding of contrastive learning. 1 2 Related Work Artifacts in Transformer Representations. Transformers are known to exhibit uneven attention allocation across input tokens. In NLP, Xiao, et al. [41] showed that early-position tokens receive disproportionate attention, regardless of their semantic importance. Sun, et al. [38] attributed such behavior to sparse, high-norm activations. Extending these observations to vision, Darcet, et al. [9] found that ViTs often produce a small set of high-norm patch tokens concentrated in background regions, which act as "registers" for global context. While these patch repurposing have been mostly studied in large-scale ViTs, our work shows it also emerge in smaller models like DINOv2-S.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖