HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
Haotian Tang, Yecheng Wu, Shang Yang, Enze Xie, Junsong Chen, Junyu Chen, Zhuoyang Zhang, Han Cai, Yao Lu, Song Han
2025Year
34Top-tier citations
Abstract
Throughput (img/s) 7.7x higher Figure 1 : HART is an early autoregressive model that can directly generate 1024×1024 images with quality comparable to diffusion models, while offering significantly improved efficiency. It achieves 4.5-7.7× higher throughput, 3.1-5.9× lower latency (measured on A100), and 6.9-13.4× lower MACs compared to state-of-the-art diffusion models. Check out our online demo and video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache CompressionKunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng HwangNeurIPS 2025 · 23 citations
- FlexVAR: Flexible Visual Autoregressive Modeling without Residual PredictionSiyu Jiao, Gengwei Zhang, Yinlong Qian, Jiancheng Huang et al.NeurIPS 2025 · 23 citations
- Soft-Masked Diffusion Language ModelsMichael Hersche, Samuel Moor-Smith, Thomas Hofmann, Abbas RahimiICLR 2026 · 15 citations
- Semantic Context Matters: Improving Conditioning for Autoregressive ModelsDongyang Jin, Ryan Xu, Jianhao Zeng, Rui Lan et al.CVPR 2026 · 12 citations
- LMM4LMM: Benchmarking and Evaluating Large-Multimodal Image Generation With LMMsJiarui Wang, Huiyu Duan, Yu Zhao, Juntong Wang et al.ICCV 2025 · 12 citations
Builds on35
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- MAGVIT: Masked Generative Video TransformerLijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama et al.CVPR 2023
- HMAR: Efficient Hierarchical Masked Auto-Regressive Image GenerationHermann Kumbong, Xian Liu, Tsung-Yi Lin, Ming-Yu Liu et al.CVPR 2025
- Sample- and Parameter-Efficient Auto-Regressive Image ModelsElad Amrani, Leonid Karlinsky, Alex M. BronsteinCVPR 2025
- Denoising Autoregressive Transformers for Scalable Text-to-Image GenerationJiatao Gu, Yuyang Wang, Yizhe Zhang, Qihang Zhang et al.ICLR 2025
- From Slow Bidirectional to Fast Autoregressive Video Diffusion ModelsTianwei Yin, Qiang Zhang, Richard Zhang, William T. Freeman et al.CVPR 2025
