Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
Joel Valdivia Ortega, Lorenz Lamm, Franziska Eckardt, Benedikt Schworm, Marion Jasnin, Tingying Peng
Abstract
Vision Transformers (ViTs), such as DINOv2, achieve strong performance across domains but often repurpose low-informative patch tokens in ways that reduce the interpretability of attention and feature maps. This challenge is especially evident in medical imaging, where domain shifts can degrade both performance and transparency. In this paper, we introduce Randomized-MLP (RMLP) regularization, a contrastive learning-based method that encourages more semantically aligned representations. We use RMLPs when fine-tuning DINOv2 to both medical and natural image modalities, showing that it improves or maintains downstream performance while producing more interpretable attention maps. We also provide a mathematical analysis of RMLPs, offering insights into its role in enhancing ViT-based models and advancing our understanding of contrastive learning. 1 2 Related Work Artifacts in Transformer Representations. Transformers are known to exhibit uneven attention allocation across input tokens. In NLP, Xiao, et al. [41] showed that early-position tokens receive disproportionate attention, regardless of their semantic importance. Sun, et al. [38] attributed such behavior to sparse, high-norm activations. Extending these observations to vision, Darcet, et al. [9] found that ViTs often produce a small set of high-norm patch tokens concentrated in background regions, which act as "registers" for global context. While these patch repurposing have been mostly studied in large-scale ViTs, our work shows it also emerge in smaller models like DINOv2-S.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba5dd322-a07b-4306-9dc5-1a465863dd0eBuilds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Register and [CLS] tokens induce a decoupling of local and global features in large ViTsAlexander Lappe, Martin A. GieseNeurIPS 2025 · 9 citations
- Random Registers for Cross-Domain Few-Shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2025
- Revisiting Continuity of Image Tokens for Cross-domain Few-shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2025
- MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationZhiwei Yang, Yucong Meng, Kexue Fu, Shuo Wang et al.AAAI 2025 · 14 citations
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 17 citations
