Semantic-Aware Human Object Interaction Image Generation
Zhu Xu, Qingchao Chen, Yuxin Peng, Yang Liu
摘要
Recent text-to-image generative models have demonstrated remarkable abilities in generating realistic images. Despite their great success, these models struggle to generate high-fidelity images with prompts oriented toward human-object interaction (HOI). The difficulty in HOI generation arises from two aspects. Firstly, the complexity and diversity of human poses challenge plausible human generation. Furthermore, untrustworthy generation of interaction boundary regions may lead to deficiency in HOI semantics. To tackle the problems, we propose a Semantic-Aware HOI generation framework SA-HOI . It utilizes human pose quality and interaction boundary region information as guidance for denoising process, thereby encouraging refinement in these regions to produce more reasonable HOI images. Based on it, we establish an iterative inversion and image refinement pipeline to continually enhance generation quality. Further, we introduce a comprehensive benchmark for HOI generation, which comprises a dataset involving diverse and fine-grained HOI categories, along with multiple custom-tailored evaluation metrics for HOI generation. Experiments demonstrate that our method significantly improves generation quality under both HOI-specific and conventional image evaluation metrics. The code is available at https://github.com/XZPKU/SA-HOI.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 被引用 78 次
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng 等ICCV 2025 · 被引用 6 次
- ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual GroundingMinghang Zheng, Jiahua Zhang, Qingchao Chen, Yuxin Peng 等ACM MM 2024 · 被引用 5 次
- InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable ObjectsXinhao Cai, Minghang Zheng, Xin Jin, Yang LiuACM MM 2025 · 被引用 1 次
- Interact-Custom: Customized Human Object Interaction Image GenerationZhu Xu, Zhaowen Wang, Yuxin Peng, Yang LiuACM MM 2025 · 被引用 1 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- EmIT: Emotional Interaction control in Text-to-image diffusion modelsHaofan Zhang, Shangfei WangACM MM 2025
- PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction GenerationXinting Hu, Haoran Wang, Jan Eric Lenssen, Bernt SchieleCVPR 2025
- Learning to Generate Human-Human-Object Interactions from Textual DescriptionsJeonghyeon Na, Sangwon Baik, Inhee Lee, Junyoung Lee 等NeurIPS 2025 · 被引用 3 次
- An Image-like Diffusion Method for Human-Object Interaction DetectionXiaofei Hui, Haoxuan Qu, Hossein Rahmani, Jun LiuCVPR 2025
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 被引用 29 次
