Customizing Text-to-Image Generation with Inverted Interaction
Mengmeng Ge, Xu Jia, Takashi Isobe, Xiaomin Li, Qinghe Wang, Jing Mu, Dong Zhou, Li Wang, Huchuan Lu, Lu Tian, Ashish Sirasao, Emad Barsoum
摘要
Subject-driven image generation, aimed at customizing user-specified subjects, has experienced rapid progress. However, most of them focus on transferring the customized appearance of subjects. In this work, we consider a novel concept customization task, that is, capturing the interaction between subjects in exemplar images and transferring the learned concept of interaction to achieve customized text-to-image generation. Intrinsically, the interaction between subjects is diverse and is difficult to describe in only a few words. In addition, typical exemplar images are about the interaction between humans, which further intensifies the challenge of interaction-driven image generation with various categories of subjects. To address this task, we adopt a divide-and-conquer strategy and propose a two-stage interaction inversion framework. The framework begins by learning a pseudo-word for a single pose of each subject in the interaction. This is then employed to promote the learning of the concept for the interaction. In addition, language prior and cross-attention loss are incorporated into the optimization process to encourage the modeling of interaction. Extensive experiments demonstrate that the proposed methods are able to effectively invert the interactive pose from exemplar images and apply it to the customized generation with user-specified interaction.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video GenerationShuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang 等CVPR 2026 · 被引用 9 次
- DreamRelation: Relation-Centric Video CustomizationYujie Wei, Shiwei Zhang, Hangjie Yuan, Biao Gong 等ICCV 2025 · 被引用 5 次
- Ego-InBetween: Generating Object State Transitions in Ego-Centric VideosMengmeng Ge, Takashi Isobe, Xu Jia, Yanan Sun 等CVPR 2026 · 被引用 1 次
- DreamRelation: Bridging Customization and Relation GenerationQingyu Shi, Lu Qi, Jianzong Wu, Jinbin Bai 等CVPR 2025
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image GenerationRuoxuan Zhang, Bin Wen, Hongxia Xie, Yi Yao 等ACM MM 2025
相关 Paper
- Decoupled Textual Embeddings for Customized Image GenerationYufei Cai, Yuxiang Wei, Zhilong Ji, Jinfeng Bai 等AAAI 2024 · 被引用 24 次
- Interact-Custom: Customized Human Object Interaction Image GenerationZhu Xu, Zhaowen Wang, Yuxin Peng, Yang LiuACM MM 2025 · 被引用 1 次
- RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image CustomizationMengqi Huang, Zhendong Mao, Mingcong Liu, Qian He 等CVPR 2024
- DynASyn: Multi-Subject Personalization Enabling Dynamic Action SynthesisYongjin Choi, Chanhun Park, Seung Jun BaekAAAI 2025 · 被引用 3 次
- AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image GenerationLianyu Pang, Jian Yin, Baoquan Zhao, Feize Wu 等NeurIPS 2024 · 被引用 18 次
