H-ViT: A Hierarchical Vision Transformer for Deformable Image Registration
Morteza Ghahremani, Mohammad Khateri, Bailiang Jian, Benedikt Wiestler, Ehsan Adeli, Christian Wachinger
Abstract
This paper introduces a novel top-down representation approach for deformable image registration, which estimates the deformation field by capturing various shortand long-range flow features at different scale levels. As a Hierarchical Vision Transformer (H-ViT), we propose a dual self-attention and cross-attention mechanism that uses high-level features in the deformation field to represent lowlevel ones, enabling information streams in the deformation field across all voxel patch embeddings irrespective of their spatial proximity. Since high-level features contain abstract flow patterns, such patterns are expected to effectively contribute to the representation of the deformation field in lower scales. When the self-attention module utilizes within-scale short-range patterns for representation, the cross-attention modules dynamically look for the key tokens across different scales to further interact with the local query voxel patches. Our method shows superior accuracy and visual quality over the state-of-the-art registration methods in five publicly available datasets, highlighting a substantial enhancement in the performance of medical imaging registration. The project link is available at https://mogvision.github.io/hvit .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus DiseasesYuxin Lin, Wei Wang, Xiaoling Luo, Zhihao Wu et al.AAAI 2025 · 3 citations
- Unsupervised Trajectory Optimization for 3D Registration in Serial Section Electron Microscopy using Neural ODEsZhenbang Zhang, Jingtong Feng, Hongjia Li, Haythem El-Messiry et al.NeurIPS 2025 · 1 citation
- PRINTER: Deformation-Aware Adversarial Learning for Virtual IHC Staining with In Situ FidelityYizhe Yuan, Bingsen Xue, Bangzheng Pu, Chengxiang Wang et al.ACM MM 2025 · 1 citation
- Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityGuangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang et al.AAAI 2026
- SACB-Net: Spatial-awareness Convolutions for Medical Image RegistrationXinxing Cheng, Tianyang Zhang, Wenqi Lu, Qingjie Meng et al.CVPR 2025
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
- CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale AttentionWenxiao Wang, Lu Yao, Long Chen, Binbin Lin et al.ICLR 2022 · 367 citations
Related papers
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He et al.ICCV 2021 · 154 citations
- Affine Medical Image Registration with Coarse-to-Fine Vision TransformerTony C. W. Mok, Albert C. S. ChungCVPR 2022 · 94 citations
- DeViT: Deformed Vision Transformers in Video InpaintingJiayin Cai, Changlin Li, Xin Tao, Chun Yuan et al.ACM MM 2022 · 15 citations
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- HiViT: A Simpler and More Efficient Design of Hierarchical Vision TransformerXiaosong Zhang, Yunjie Tian, Lingxi Xie, Wei Huang et al.ICLR 2023
