Interact-Custom: Customized Human Object Interaction Image Generation
Zhu Xu, Zhaowen Wang, Yuxin Peng, Yang Liu
摘要
Compositional Customized Image Generation aims to customize multiple target concepts within generation content, which has gained attention for its wild application. Though a great success, existing approaches mainly concentrate on the target entity's appearance preservation, while neglecting the fine-grained interaction control among target entities. To enable the model of such interaction control capability, we focus on human object interaction scenario and propose the task of Customized Human Object Interaction Image Generation (CHOI), which simultaneously requires identity preservation for target human object and the interaction semantic control between them. We attribute two primary challenges of CHOI as follows: (1) the simultaneous identity preservation and interaction control demands require the model to decompose the human object into self-contained identity features and pose-oriented interaction features, while the current HOI image datasets fail to provide ideal samples for such feature-decomposed learning. (2) inappropriate spatial configuration between human and object may lead to the lack of desired interaction semantics, as it may provide wrong hints on the human object body parts crucial for interaction semantic expression. To tackle the above issues, we first collect and process a large-scale dataset, where each sample encompasses the same pair of human object involving different interactive poses. Such data is tailored for CHOI training, from where the model can learn how to decompose identity features and interaction features for target human and object. Then to provide appropriate spatial configuration for interaction semantic expression, we design a two-stage model Interact-Custom, which firstly explicitly model the spatial configuration by generating a foreground mask depicting the interaction behavior, then under the guidance of this mask, we generate the target human object interacting while preserving their identities features. Furthermore, if the background image and the union location of where the target human object should appear are provided by users, Interact-Custom also provides the optional functionality to specify them, offering high content controllability. Extensive experiments on our tailored metrics for CHOI task demonstrate the effectiveness of our approach. Our code is available at https://github.com/XZPKU/Inter-custom.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper24
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik 等ICLR 2023 · 被引用 464 次
- Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion ModelsYuchao Gu, Xintao Wang, Jay Zhangjie Wu, Yujun Shi 等NeurIPS 2023 · 被引用 333 次
- Key-Locked Rank One Editing for Text-to-Image PersonalizationYoad Tewel, Rinon Gal, Gal Chechik, Yuval AtzmonSIGGRAPH 2023 · 被引用 119 次
相关 Paper
- Customizing Text-to-Image Generation with Inverted InteractionMengmeng Ge, Xu Jia, Takashi Isobe, Xiaomin Li 等ACM MM 2024 · 被引用 3 次
- HOComp: Interaction-Aware Human-Object CompositionDong Liang, Jinyuan Jia, Yuhao Liu, Rynson W. H. LauNeurIPS 2025 · 被引用 1 次
- Semantic-Aware Human Object Interaction Image GenerationZhu Xu, Qingchao Chen, Yuxin Peng, Yang LiuICML 2024 · 被引用 9 次
- PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction GenerationXinting Hu, Haoran Wang, Jan Eric Lenssen, Bernt SchieleCVPR 2025
- EmIT: Emotional Interaction control in Text-to-image diffusion modelsHaofan Zhang, Shangfei WangACM MM 2025
