MagicFight: Personalized Martial Arts Combat Video Generation
Jiancheng Huang, Mingfu Yan, Songyan Chen, Yi Huang, Shifeng Chen
Abstract
Amid the surge in generic text-to-video generation, the field of personalized human video generation has witnessed notable advancements, primarily concentrated on single-person scenarios. However, to our knowledge, the domain of two-person interactions, particularly in the context of martial arts combat, remains uncharted. We identify a significant gap: existing models for single-person dancing generation prove insufficient for capturing the subtleties and complexities of two engaged fighters, resulting in challenges such as identity confusion, anomalous limbs, and action mismatches. To address this, we introduce a pioneering new task, Personalized Martial Arts Combat Video Generation. Our approach, MagicFight, is specifically crafted to overcome these hurdles. Given this pioneering task, we face a lack of appropriate datasets. Thus, we generate a bespoke dataset using the game physics engine Unity, meticulously crafting a multitude of 3D characters, martial arts moves, and scenes designed to represent the diversity of combat. MagicFight refines and adapts existing models and strategies to generate high-fidelity two-person combat videos that maintain individual identities and ensure seamless, coherent action sequences, thereby laying the groundwork for future innovations in the realm of interactive video content creation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingShuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu et al.NeurIPS 2025 · 228 citations
- Freehand Sketch Generation from Mechanical ComponentsZhichao Liao, Fengyuan Piao, Di Huang, Xinghui Li et al.ACM MM 2024 · 12 citations
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyQuanjian Song, Donghao Zhou, Jingyu Lin, Fei Shen et al.NeurIPS 2025 · 9 citations
- OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video GenerationDonghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li et al.ICML 2026 · 3 citations
- GameGen-X: Interactive Open-world Game Video GenerationHaoxuan Che, Xuanhua He, Quande Liu, Cheng Jin et al.ICLR 2025 · 2 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Magicid: Hybrid Preference Optimization for Id-Consistent and Dynamic-Preserved Video CustomizationHengjia Li, Lifan Jiang, Xi Xiao, Tianyang Wang et al.ICCV 2025 · 2 citations
- Multi-Identity Human Image Animation with Structural Video DiffusionZhenzhi Wang, Yixuan Li, Yanhong Zeng, Yuwei Guo et al.ICCV 2025 · 1 citation
- VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion ModelsChi-Pin Huang, Yen-Siang Wu, Hung-Kai Chung, Kai-Po Chang et al.CVPR 2025
- HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language ModelsXiao Wang, Jingyun Hua, Weihong Lin, Yuanxing Zhang et al.ACL 2025 · 1 citation
- MotionCharacter: Fine-Grained Motion Controllable Human Video GenerationHaopeng Fang, Di Qiu, Binjie Mao, He TangAAAI 2026
