FiT: Flexible Vision Transformer for Diffusion Model
Zeyu Lu, Zidong Wang, Di Huang, Chengyue Wu, Xihui Liu, Wanli Ouyang, Lei Bai
摘要
Nature is infinitely resolution-free. In the context of this reality, existing diffusion models, such as Diffusion Transformers, often face challenges when processing image resolutions outside of their trained domain. To overcome this limitation, we present the Flexible Vision Transformer (FiT), a transformer architecture specifically designed for generating images with unrestricted resolutions and aspect ratios. Unlike traditional methods that perceive images as static-resolution grids, FiT conceptualizes images as sequences of dynamically-sized tokens. This perspective enables a flexible training strategy that effortlessly adapts to diverse aspect ratios during both training and inference phases, thus promoting resolution generalization and eliminating biases induced by image cropping. Enhanced by a meticulously adjusted network structure and the integration of training-free extrapolation techniques, FiT exhibits remarkable flexibility in resolution extrapolation generation. Comprehensive experiments demonstrate the exceptional performance of FiT across a broad range of resolutions, showcasing its effectiveness both within and beyond its training resolution distribution. Repository available at https://github.com/whlzy/FiT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiTLe Zhuo, Ruoyi Du, Han Xiao, Yangguang Li 等NeurIPS 2024 · 被引用 144 次
- DDT: Decoupled Diffusion TransformerShuai Wang, Zhi Tian, Weilin Huang, Limin WangCVPR 2026 · 被引用 102 次
- Mitigating Object Hallucination via Concentric Causal AttentionYun Xing, Yiheng Li, Ivan Laptev, Shijian LuNeurIPS 2024 · 被引用 78 次
- PixNerd: Pixel Neural Field DiffusionShuai Wang, Ziteng Gao, Chenhui Zhu, Weilin Huang 等ICLR 2026 · 被引用 78 次
- U-DiTs: Downsample Tokens in U-Shaped Diffusion TransformersYuchuan Tian, Zhijun Tu, Hanting Chen, Jie Hu 等NeurIPS 2024 · 被引用 59 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- ResAdapter: Domain Consistent Resolution Adapter for Diffusion ModelsJiaxiang Cheng, Pan Xie, Xin Xia, Jiashi Li 等AAAI 2025 · 被引用 3 次
- Native-Resolution Image SynthesisZidong Wang, Lei Bai, Xiangyu Yue, Wanli Ouyang 等NeurIPS 2025 · 被引用 14 次
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution DiffusionNoam Issachar, Guy Yariv, Sagie Benaim, Yossi Adi 等ICML 2026 · 被引用 14 次
- All are Worth Words: A ViT Backbone for Diffusion ModelsFan Bao, Shen Nie, Kaiwen Xue, Yue Cao 等CVPR 2023
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and ResolutionMostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek 等NeurIPS 2023 · 被引用 303 次
