Revisiting Model Stitching In the Foundation Model Era
Zheda Mai, Ke Zhang, Fu-En Wang, Zixiao Ken Wang, Albert Y. C. Chen, Lu Xia, Min Sun, Wei-Lun Chao, Cheng-Hao Kuo
Abstract
Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility. Prior work finds that models trained on the same dataset remain stitchable (negligible accuracy drop) despite different initializations or objectives. We revisit stitching for Vision Foundation Models (VFMs) that vary in objectives, data, and modality (e.g., CLIP, DINOv2, SigLIP2) and ask: Are heterogeneous VFMs stitchable? We introduce a systematic protocol spanning stitch positions, stitch layer families, training losses, and downstream tasks. Three findings emerge. (1) Stitch layer training matters: conventional approaches that match the intermediate features at the stitch position or optimize the task loss end-to-end struggle to retain accuracy, especially at shallow stitch positions. (2) With a simple feature-matching loss at the target model's penultimate layer, heterogeneous VFMs become reliably stitchable across vision tasks. (3) For deep stitch positions, the stitched model can significantly surpass either constituent model with a small inference overhead (for the stitch layer). Building on these findings, we further propose the VFM Stitch Tree (VST), which shares early layers across VFMs while retaining their later layers, yielding a controllable accuracy-latency trade-off for multimodal LLMs that often leverage multiple VFMs. Taken together, our study elevates stitching from a diagnostic probe to a practical recipe for integrating complementary VFM strengths and pinpointing where their representations align or diverge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 731e729f-14f0-4a68-8b33-df69f97948c5Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation ModelsSofian Chaybouti, Sanath Narayan, Yasser Dahou, Phúc H. Lê Khắc et al.CVPR 2026 · 2 citations
- AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into OneMike Ranzinger, Greg Heinrich, Jan Kautz, Pavlo MolchanovCVPR 2024 · 31 citations
- Latent Space Translation via Semantic AlignmentValentino Maiorca, Luca Moschella, Antonio Norelli, Marco Fumero et al.NeurIPS 2023 · 59 citations
- Functional Alignment Can Mislead: Examining Model StitchingDamian Smith, Harvey Mannering, Antonia MarcuICML 2025
- All-in-One: Transferring Vision Foundation Models into Stereo MatchingJingyi Zhou, Haoyu Zhang, Jiakang Yuan, Peng Ye et al.AAAI 2025
