Action Imitation in Common Action Space for Customized Action Image Synthesis
Wang Lin, Jingyuan Chen, Jiaxin Shi, Zirun Guo, Yichen Zhu, Zehan Wang, Tao Jin, Zhou Zhao, Fei Wu, Shuicheng Yan, Hanwang Zhang
Abstract
We propose a novel method, TwinAct , to tackle the challenge of decoupling actions and actors in order to customize the text-guided diffusion models (TGDMs) for few-shot action image generation. TwinAct addresses the limitations of existing methods that struggle to decouple actions from other semantics ( e.g. , the actor’s appearance) due to the lack of an effective inductive bias with few exemplar images. Our approach introduces a common action space, which is a textual embedding space focused solely on actions, enabling precise customization without actor-related details. Specifically, TwinAct involves three key steps: 1) Building common action space based on a set of representative action phrases; 2) Imitating the customized action within the action space; and 3) Generating highly adaptable customized action images in diverse contexts with action similarity loss. To comprehensively evaluate TwinAct, we construct a novel benchmark, which provides sample images with various forms of actions. Extensive experiments demonstrate TwinAct’s superiority in generating accurate, context-independent customized actions while maintaining the identity consistency of different subjects, including animals, humans, and even customized actors. Project page: https://twinact-official.github.io/TwinAct/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a944f8e-ea0e-4052-9c11-a216e30ebf9dCited by top-tier papers9
- Knowledge Is Power: Harnessing Large Language Models for Enhanced Cognitive DiagnosisZhiang Dong, Jingyuan Chen, Fei WuAAAI 2025 · 15 citations
- WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed BenchmarkWang Lin, Feng Wang, Majun Zhang, Wentao Hu et al.ICLR 2026 · 2 citations
- Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across DomainsFangming Feng, Sihang Cai, Zequn Xie, Yangyang Wu et al.AAAI 2026 · 1 citation
- ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion MitigationZirun Guo, Tao JinCVPR 2025
- Vinci: Deep Thinking in Text-to-Image Generation using Unified Model with Reinforcement LearningWang Lin, Wentao Hu, Liyu Jia, Kaihang Pan et al.NeurIPS 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image CustomizationYeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin et al.AAAI 2025 · 6 citations
- DynASyn: Multi-Subject Personalization Enabling Dynamic Action SynthesisYongjin Choi, Chanhun Park, Seung Jun BaekAAAI 2025 · 3 citations
- OneActor: Consistent Subject Generation via Cluster-Conditioned GuidanceJiahao Wang, Caixia Yan, Haonan Lin, Weizhan Zhang et al.NeurIPS 2024 · 16 citations
- Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image PersonalizationHenglei Lv, Jiayu Xiao, Liang LiACM MM 2024 · 6 citations
- ACT: A Unified Framework for Rigging and Animating Characters with Arbitrary TopologiesPengyu Long, Weirui Wang, Qingcheng Zhao, Xiaoyang Guo et al.SIGGRAPH 2026
