FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
Shiyi Zhang, Junhao Zhuang, Zhaoyang Zhang, Ying Shan, Yansong Tang
Abstract
Action customization involves generating videos where the subject performs actions dictated by input control signals. Current methods use pose-guided or global motion customization but are limited by strict constraints on spatial structure such as layout, skeleton, and viewpoint consistency, reducing adaptability across diverse subjects and scenarios. To overcome these limitations, we propose FlexiAct, which transfers actions from a reference video to an arbitrary target image. Unlike existing methods, FlexiAct allows for variations in layout, viewpoint, and skeletal structure between the subject of the reference video and the target image, while maintaining identity consistency. Achieving this requires precise action control, spatial structure adaptation, and consistency preservation. To this end, we introduce RefAdapter, a lightweight image-conditioned adapter that excels in spatial adaptation and consistency preservation, surpassing existing methods in balancing appearance consistency and structural flexibility. Additionally, based on our observations, the denoising process exhibits varying levels of attention to motion (low frequency) and appearance details (high frequency) at different timesteps. So we propose FAE (Frequency-aware Action Extraction), which, unlike existing methods that rely on separate spatial-temporal architectures, directly achieves action extraction during the denoising process. Experiments demonstrate that our method effectively transfers actions to subjects with diverse layouts, skeletons, and viewpoints. We release our code and model weights to support further research at FlexiAct.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96e04059-a4d7-40c7-8af1-b8182626c803Cited by top-tier papers12
- Video-As-Prompt: Unified Semantic Control for Video GenerationYuxuan Bian, Xin Chen, Zenan Li, Tiancheng Zhi et al.ICLR 2026 · 13 citations
- SMRABooth: Subject and Motion Representation Alignment for Customized Video GenerationXuancheng Xu, Yaning Li, Sisi You, Bing-Kun BaoCVPR 2026 · 11 citations
- Frame In-N-Out: Unbounded Controllable Image-to-Video GenerationBoyang Wang, Xuweiyi Chen, Matheus Gadelha, Zezhou ChengNeurIPS 2025 · 9 citations
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video GenerationShuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang et al.CVPR 2026 · 9 citations
- TS-Attn: Temporal-wise Separable Attention for Multi-Event Video GenerationHongyu Zhang, Yufan Deng, Zilin Pan, Peng-Tao Jiang et al.ICLR 2026 · 5 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion TransferLi Yuze, Dong Gong, Xiao Cao, Junchao Yuan et al.CVPR 2026 · 3 citations
- Zero-to-Hero: Empowering Video Appearance Transfer with Zero-Shot Initialization and Holistic RestorationTongtong Su, Chengyu Wang, Haipeng Liao, Jun Huang et al.AAAI 2026
- FlexNeRF: Photorealistic Free-viewpoint Rendering of Moving Humans from Sparse ViewsVinoj Jayasundara, Amit Agrawal, Nicolas Heron, Abhinav Shrivastava et al.CVPR 2023
- MotionEditor: Editing Video Motion via Content-Aware DiffusionShuyuan Tu, Qi Dai, Zhi-Qi Cheng, Han Hu et al.CVPR 2024 · 21 citations
- MetaPix: Few-Shot Video RetargetingJessica Lee, Deva Ramanan, Rohit GirdharICLR 2020 · 29 citations
