Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
Joonhyung Park, Hyeongwon Jang, Joowon Kim, Eunho Yang
Abstract
Recent visual autoregressive (AR) models have shown promising capabilities in text-to-image generation, operating in a manner similar to large language models. While test-time computation scaling has brought remarkable success in enabling reasoning-enhanced outputs for challenging natural language tasks, its adaptation to visual AR models remains unexplored and poses unique challenges. Naively applying test-time scaling strategies such as Best-of-N can be suboptimal: they consume full-length computation on erroneous generation trajectories, while the raster-scan decoding scheme lacks a blueprint of the entire canvas, limiting scaling benefits as only a few prompt-aligned candidates are generated. To address these, we introduce GridAR, a test-time scaling framework designed to elicit the best possible results from visual AR models. GridAR employs a grid-partitioned progressive generation scheme in which multiple partial candidates for the same position are generated within a canvas, infeasible ones are pruned early, and viable ones are fixed as anchors to guide subsequent decoding. Coupled with this, we present a layout-specified prompt reformulation strategy that inspects partial views to infer a feasible layout for satisfying the prompt. The reformulated prompt then guides subsequent image generation to mitigate the blueprint deficiency. Together, GridAR achieves higher-quality results under limited test-time scaling: with N=4, it even outperforms Best-of-N (N=8) by 14.4% on T2I-CompBench++ while reducing cost by 25.6%. It also generalizes to autoregressive image editing, showing comparable edit quality and a 13.9% gain in semantic preservation on PIE-Bench over larger-N baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7084b64-b35a-469a-aa43-c06bb02edbb0Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- ScalingAR: Scaling Confidence for Autoregressive Image GenerationHarold Haodong Chen, Xianfeng Wu, Wenjie Shu, Rongjin Guo et al.ICML 2026 · 3 citations
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive GenerationZhekai Chen, Ruihang Chu, Yukang Chen, Shiwei Zhang et al.NeurIPS 2025 · 16 citations
- Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual GenerationSubin Kim, Sangwoo Mo, Mamshad Nayeem Rizve, Yiran Xu et al.CVPR 2026 · 2 citations
- Neighboring Autoregressive Modeling for Efficient Visual GenerationYefei He, Yuanyu He, Shaoxuan He, Feng Chen et al.ICCV 2025 · 3 citations
- VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning ModelsSoumya Suvra Ghosal, Youngeun Kim, Zhuowei Li, Ritwick Chaudhry et al.CVPR 2026
