SCALAR: Scale-wise Controllable Visual Autoregressive Learning
Ryan Xu, Dongyang Jin, Yancheng Bai, Rui Lan, Xu Duan, Lei Sun, Xiangxiang Chu
摘要
Controllable image synthesis, which enables fine-grained control over generated outputs, has emerged as a key focus in visual generative modeling. However, controllable generation remains challenging for Visual Autoregressive (VAR) models due to their hierarchical, next-scale prediction style. Existing VAR-based methods often suffer from inefficient control encoding and disruptive injection mechanisms that compromise both fidelity and efficiency. In this work, we present SCALAR, a controllable generation method based on VAR, incorporating a Scale-wise Conditional Decoding mechanism. SCALAR leverages a pretrained image encoder to extract semantic control signal encodings, which are projected into scale-specific representations and injected into the corresponding layers of the VAR backbone. This design provides persistent and structurally aligned guidance throughout the generation process. Building on SCALAR, we develop SCALAR-Uni, a unified extension that aligns multiple control modalities into a shared latent space, supporting flexible multi-conditional guidance in a single model. Extensive experiments show that SCALAR achieves superior generation quality and control precision across various tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement LearningChubin Chen, Sujie Hu, Jiashu Zhu, Meiqi Wu 等CVPR 2026 · 被引用 28 次
- Semantic Context Matters: Improving Conditioning for Autoregressive ModelsDongyang Jin, Ryan Xu, Jianhao Zeng, Rui Lan 等CVPR 2026 · 被引用 12 次
- From Scale to Speed: Adaptive Test-Time Scaling for Image EditingXiangyan Qu, Zhenlong Yuan, Jing Tang, Rui Chen 等CVPR 2026 · 被引用 8 次
- Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive ModelsYexing Xu, Wei Feng, Shen Zhang, Haohan Wang 等CVPR 2026 · 被引用 1 次
- AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured ReasoningBinxiao Xu, Junyu Feng, Xiaopeng Lin, Haodong Li 等ICML 2026
它引用的顶会 Paper30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- FlowAR: Scale-wise Autoregressive Image Generation Meets Flow MatchingSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen 等ICML 2025
- Collaborative Decoding Makes Visual Auto-Regressive Modeling EfficientZigeng Chen, Xinyin Ma, Gongfan Fang, Xinchao WangCVPR 2025
- Visual Autoregressive Modeling for Instruction-Guided Image EditingQingyang Mao, Qi Cai, Yehao Li, Yingwei Pan 等ICLR 2026 · 被引用 21 次
- ControlAR: Controllable Image Generation with Autoregressive ModelsZongming Li, Tianheng Cheng, Shoufa Chen, Peize Sun 等ICLR 2025
- Visual Autoregressive Modeling for Image Super-ResolutionYunpeng Qu, Kun Yuan, Jinhua Hao, Kai Zhao 等ICML 2025
