ViT^3: Unlocking Test-Time Training in Vision
Dongchen Han, Yining Li, Tianyu Li, Zixuan Cao, Ziming Wang, Jun Song, Cheng Yu, Bo Zheng, Gao Huang
Abstract
Test-Time Training (TTT) has recently emerged as a promising direction for efficient sequence modeling. TTT reformulates attention operation as an online learning problem, constructing a compact inner model from keyvalue pairs at test time. This reformulation opens a rich and flexible design space while achieving linear computational complexity. However, crafting a powerful visual TTT design remains challenging: fundamental choices for the inner module and inner training lack comprehensive understanding and practical guidelines. To bridge this critical gap, in this paper, we present a systematic empirical study of TTT designs for visual sequence modeling. From a series of experiments and analyses, we distill six practical insights that establish design principles for effective visual TTT and illuminate paths for future improvement. These findings culminate in the Vision Test-Time Training (ViT 3 ) model, a pure TTT architecture that achieves linear complexity and parallelizable computation. We evaluate ViT 3 across diverse visual tasks, including image classification, image generation, object detection, and semantic segmentation. Results show that ViT 3 consistently matches or outperforms advanced linear-complexity models (e.g., Mamba and linear attention variants) and effectively narrows the gap to highly optimized vision Transformers. We hope this study and the ViT 3 baseline can facilitate future work on visual TTT models. Code: github.com/LeapLabTHU/ViTTT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8355adb-6133-41d3-ad1b-3d2ad3db4c66Cited by top-tier papers4
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single PassYewei Liu, Xiyuan Wang, Yansheng Mao, Yoav Gelberg et al.ICML 2026 · 11 citations
- Test-Time Training with KV Binding Is Secretly Linear AttentionJunchen Liu, Sven Elflein, Or Litany, Zan Gojcic et al.ICML 2026 · 6 citations
- Linearizing Vision Transformer with Test-Time TrainingYining Li, Dongchen Han, Zeyu Liu, Hanyi Wang et al.ICML 2026 · 1 citation
- VGG-T: Offline Feed-Forward 3D Reconstruction at ScaleSven Elflein, Ruilong Li, Sérgio Agostinho, Zan Gojcic et al.CVPR 2026
Builds on48
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- Learning to (Learn at Test Time): RNNs with Expressive Hidden StatesYu Sun, Xinhao Li, Karan Dalal, Jiarui Xu et al.ICML 2025
- Test-Time Training Done RightTianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang et al.ICLR 2026 · 127 citations
- PRISM: Parallel Residual Iterative Sequence ModelJie Jiang, Ke Cheng, XIN XU, Mengyang Pang et al.ICML 2026
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and ResolutionMostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek et al.NeurIPS 2023 · 303 citations
- ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision ModelsGuoyizhe Wei, Rama ChellappaICCV 2025
