StructMAR: Structure-Aware Masked Autoregression for Explicit Layout Alignment in Text-to-Image Generation
Gang Cao, Junying Zhang
摘要
Although text-to-image generation has achieved significant progress, strict instance-level layout alignment remains a challenge for many applications. Masked Autoregressive (MAR) models on continuous latents are both efficient and high-fidelity, yet the standard practice of flattening 2D latents into 1D sequences weakens spatial topology, limiting precise controllability. To address this, we propose StructMAR, a structure-aware masked autoregressive framework that transforms layout alignment from a soft correlation into an explicit structural alignment. By integrating 2D Rotary Positional Embeddings with a Layout-Guided Attention Bias, StructMAR explicitly biases latent tokens toward their corresponding layout instances during attention computation. We further use Group Relative Policy Optimization (GRPO) as a final-stage policy refinement to reduce the mismatch between the MAR training objective and detector-based layout evaluation metrics. Evaluated on the COCO-Position and COCO-MIG benchmarks, StructMAR achieves state-of-the-art performance, reaching 57.2 AP and 79.4 mIoU on the former, and 61.7 ISR and 56.9 mIoU on the latter. These results, coupled with a 4.05 inference speedup, underscore the efficacy of explicit structural inductive biases in controllable autoregressive generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov 等ICLR 2024 · 被引用 816 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot 等ICML 2023 · 被引用 751 次
相关 Paper
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image GenerationYifu Luo, Xinhao Hu, Keyu Fan, Haoyuan Sun 等NeurIPS 2025 · 被引用 12 次
- GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment TuningYaqi Li, Han Fang, Zerun Feng, Kaijing Ma 等ACM MM 2024 · 被引用 1 次
- Seeing What Matters: Visual Preference Policy Optimization for Visual GenerationZiqi Ni, Yuanzhi Liang, Rui Li, Yi Zhou 等CVPR 2026 · 被引用 9 次
- Position-LoRA: Enhanced Relation Customization through Structural Prior in Initial Latent NoiseYiming Li, Peng Zhou, Xiaokang Qin, Hongwei Hu 等ACM MM 2025
- ConsistCompose: Unified Multimodal Layout Control for Image CompositionXuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai 等CVPR 2026 · 被引用 5 次
