HilbertA: Hilbert-Curve–Aligned Sparse Attention for 2D Structured Data
Shaoyi Zheng, Wenbo Lu, Yuxuan Xia, Shenji Wan
Abstract
Designing sparse attention for 2D image data in diffusion and vision-language models requires reconciling spatial locality with hardware-efficient execution: handcrafted 2D sparsity patterns preserve spatial structure but often induce uncoalesced memory access, limiting practical speedups on modern GPUs. We present HilbertA , a 2D-aware sparse attention mechanism that reorders image tokens along a Hilbert curve, converting local spatial neighborhoods into contiguous memory segments for efficient GPU execution. To enable communication beyond local tiles, HilbertA shifts attention windows along the Hilbert-ordered sequence across layers and uses a small central shared region, preserving contiguous access while supporting cross-tile information flow. Across diffusion and vision-language models, HilbertA delivers consistent efficiency gains while maintaining competitive quality, achieving up to 4.16× attention acceleration and 1.44× end-to-end speedup on Flux.1-dev, and up to 2.30× attention acceleration with 1.57× faster time-to-first-token on Qwen3-VL-8B inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang et al.NeurIPS 2024 · 1,029 citations
Related papers
- Hilbert-Guided Sparse Local AttentionYunge Li, Lanyu XuICLR 2026 · 1 citation
- VORTA: Efficient Video Diffusion via Routing Sparse AttentionWenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao et al.NeurIPS 2025 · 25 citations
- DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video GenerationJie Hu, Zixiang Gao, Yutong He, Kun YuanICML 2026
- Trainable Log-linear Sparse Attention for Efficient Diffusion TransformersYifan Zhou, Zeqi Xiao, Tianyi Wei, Shuai Yang et al.CVPR 2026 · 6 citations
- Faster Video Diffusion with Trainable Sparse AttentionPeiyuan Zhang, Yongqi Chen, Haofeng Huang, Will Lin et al.NeurIPS 2025 · 6 citations
