Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?
Yujin Han, Andi Han, Wei Huang, Chaochao Lu, Difan Zou
摘要
Despite the remarkable success of diffusion models (DMs) in data generation, they exhibit specific failure cases with unsatisfactory outputs. We focus on one such limitation: the ability of DMs to learn hidden rules between image features. Specifically, for image data with dependent features (x) and (y) (e.g., the height of the sun (x) and the length of the shadow (y)), we investigate whether DMs can accurately capture the inter-feature rule (p(y|x)). Empirical evaluations on mainstream DMs (e.g., Stable Diffusion 3.5) reveal consistent failures, such as inconsistent lighting-shadow relationships and mismatched object-mirror reflections. Inspired by these findings, we design four synthetic tasks with strongly correlated features to assess DMs' rule-learning abilities. Extensive experiments show that while DMs can identify coarse-grained rules, they struggle with fine-grained ones. Our theoretical analysis demonstrates that DMs trained via denoising score matching (DSM) exhibit constant errors in learning hidden rules, as the DSM objective is not compatible with rule conformity. To mitigate this, we introduce a common technique -incorporating additional classifier guidance during sampling, which achieves (limited) improvements. Our analysis reveals that the subtle signals of fine-grained rules are challenging for the classifier to capture, providing insights for future exploration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Understanding Representation Dynamics of Diffusion Models via Low-Dimensional ModelingXiao Li, Zekai Zhang, Xiang Li, Siyi Chen 等NeurIPS 2025 · 被引用 19 次
- Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMsYujin Han, Hao Chen, Andi Han, Zhiheng Wang 等ICLR 2026 · 被引用 9 次
- Personalized Federated Training of Diffusion Models with Privacy GuaranteesKumar Kshitij Patel, Bingqing Jiang, A. F. M. Mahfuzul Kabir, Weitong Zhang 等CVPR 2026
- Evaluating the Representation Space of Diffusion Models via Self-Supervised PrinciplesXiao Li, Yixuan Jia, Zekai Zhang, Xiang Li 等ICML 2026
- Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video ReconstructionYujie Wei, Chenglong Ma, Jianxiong Gao, Chenhui Wang 等CVPR 2026
它引用的顶会 Paper51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Reflected Diffusion ModelsAaron Lou, Stefano ErmonICML 2023 · 被引用 84 次
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera 等NeurIPS 2023 · 被引用 371 次
- What the DAAM: Interpreting Stable Diffusion Using Cross AttentionRaphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang 等ACL 2023 · 被引用 93 次
- Consistent Diffusion Models: Mitigating Sampling Drift by Learning to be ConsistentGiannis Daras, Yuval Dagan, Alex Dimakis, Constantinos DaskalakisNeurIPS 2023 · 被引用 79 次
- Denoising Diffusion Bridge ModelsLinqi Zhou, Aaron Lou, Samar Khanna, Stefano ErmonICLR 2024 · 被引用 163 次
