Mitigating Noise-Induced Layout Priors for Object Counting in Diffusion Models
Xiaoling Gu, Xuelong Li, Shengqi Wu, Yongkang Wong, wu, Huan Li, Zhou Yu, Mohan Kankanhalli
Abstract
Despite remarkable progress in text-to-image diffusion models, accurately generating the specified number of objects remains a persistent challenge. We identify the initial noise as a primary determinant of spatial layout formation, with early-stage cross-attention serving as the key mechanism that mediates the propagation of noise-induced structures throughout the denoising process. We characterize this phenomenon as Noise-Induced Layout Prior. Leveraging this insight, we propose a novel training-free framework for object counting in diffusion models. Our approach consists of two key components: (1) a Count-Aware Noise Adjustment Strategy, which explicitly manipulates the initial latent noise to align layout formation with the target object count, and (2) an Attention-Guided Layout Consistency Strategy, which performs test-time optimization on early-stage cross-attention to further stabilize layout formation during denoising. Extensive experiments on both single-category and multi-category benchmarks demonstrate that our method consistently outperforms strong diffusion baselines and state-of-the-art object count control methods in terms of counting accuracy and image quality. Code Release: [https://github.com/lxlong1201/Mitigate Noise Prior](https://github.com/lxlong1201/Mitigate Noise Prior).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45e62bf9-a047-40a5-9ba5-52588a3c7bb4Builds on18
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Be Decisive: Noise-Induced Layouts for Multi-Subject GenerationOmer Dahary, Yehonathan Cohen, Or Patashnik, Kfir Aberman et al.SIGGRAPH 2025 · 3 citations
- Control and Realism: Best of Both Worlds in Layout-to-Image without TrainingBonan Li, Yinhan Hu, Songhua Liu, Xinchao WangICML 2025
- Dense Text-to-Image Generation with Attention ModulationYunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha et al.ICCV 2023 · 204 citations
- LoCo: Training-Free Layout-to-Image Synthesis with Localized ConstraintsPeiang Zhao, Han Li, Ruiyang Jin, S. Kevin ZhouACM MM 2025 · 2 citations
- WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single ImageJiwoo Park, Tae Eun Choi, Youngjun Jun, Seong Jae HwangICCV 2025
