A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image Generation
Shuang Hao, Pengfei Ren, Haifeng Sun, Pan Ting, Qi Qi, Lei Zhang, Cong Liu, Jianxin Liao, Jingyu Wang
摘要
Controllable hand image generation aims to synthesize geometrically accurate images with consistent appearance. Recently, diffusion models have been widely applied for hand image synthesis. However, through input-level fusion or feature-level modulation, existing methods inject control signals with fixed strength across all timesteps, ignoring the progressive nature of the denoising process. In this paper, we reveal that the modulation of control signals depends on the denoising state and condition complexity. Due to distinct semantic distributions and information densities, achieving effective interaction among these heterogeneous representations remains a challenge. To address this, we propose a Temporal and Content Co-Awareness Latent Diffusion method that introduces a dual-driven modulation strategy. Specifically, we design a query-based interaction mechanism to mitigate information redundancy and align semantic distributions. Leveraging cross-domain interaction, the model infers required control information to dynamically adjust pose and appearance injection strengths. Furthermore, we design a Pose-Invariant Appearance Encoder that captures both global appearance consistency and local texture details. Extensive experiments validate our superiority over state-of-the-art. Code is available at https://github.com/samukahs/TCCA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano 等CVPR 2022 · 被引用 984 次
- FSGAN: Subject Agnostic Face Swapping and ReenactmentYuval Nirkin, Yosi Keller, Tal HassnerICCV 2019 · 被引用 710 次
相关 Paper
- Controllable Person Image Synthesis with Pose-Constrained Latent DiffusionXiao Han, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song 等ICCV 2023 · 被引用 36 次
- Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisYanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie 等CVPR 2024 · 被引用 26 次
- Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image DiffusionJiwon Kim, Pu-Reum Kim, Seonhwa Kim, Soobin Park 等ICCV 2025
- IntrinsicControlNet: Cross-Distribution Image Generation with Real and UnrealJiayuan Lu, Rengan Xie, Zixuan Xie, Zhizhen Wu 等ICCV 2025 · 被引用 4 次
- A Dual-Branch 3D Spatial-Aware Latent Diffusion for Realistic Depth Image SynthesisShuang Hao, Pengfei Ren, Lei Zhang, Haifeng Sun 等ACM MM 2025
