Rethink Sparse Signals for Pose-Guided Text-to-Image Generation
Wenjie Xuan, Jing Zhang, Juhua Liu, Bo Du, Dacheng Tao
摘要
Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-image generation. However, dense representations raised new challenges, including editing difficulties and potential inconsistencies with textual prompts. This fact motivates us to revisit sparse signals for pose guidance, owing to their simplicity and shape-agnostic nature, which remains underexplored. This paper proposes a novel Spatial-Pose ControlNet(SP-Ctrl), equipping sparse signals with robust controllability for pose-guided image generation. Specifically, we extend OpenPose to a learnable spatial representation, making keypoint embeddings discriminative and expressive. Additionally, we introduce keypoint concept learning, which encourages keypoint tokens to attend to the spatial positions of each keypoint, thus improving pose alignment. Experiments on animal- and human-centric image generation tasks demonstrate that our method outperforms recent spatially controllable T2I generation approaches under sparse-pose guidance and even matches the performance of dense signal-based methods. Moreover, SP-Ctrl shows promising capabilities in diverse and cross-species generation through sparse signals. Codes will be available at https://github.com/DREAMXFAR/SP-Ctrl.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu 等AAAI 2024 · 被引用 1,641 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 等NeurIPS 2023 · 被引用 709 次
相关 Paper
- DisPose: Disentangling Pose Guidance for Controllable Human Image AnimationHongxiang Li, Yaowei Li, Yuhang Yang, Junjie Cao 等ICLR 2025
- Readout Guidance: Learning Control from Diffusion FeaturesGrace Luo, Trevor Darrell, Oliver Wang, Dan B. Goldman 等CVPR 2024 · 被引用 11 次
- CPO: Condition Preference Optimization for Controllable Image GenerationZonglin Lyu, Ming Li, Xinxin Liu, Chen ChenNeurIPS 2025 · 被引用 1 次
- Controllable Person Image Synthesis with Pose-Constrained Latent DiffusionXiao Han, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song 等ICCV 2023 · 被引用 36 次
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free VideosYue Ma, Yingqing He, Xiaodong Cun, Xintao Wang 等AAAI 2024 · 被引用 318 次
