Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative Modeling
Chenghao Xu, Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng
摘要
In dance performances, choreographers define the visual expression of movement, while cinematographers shape its final presentation through camera work. Consequently, the synthesis of camera movements informed by both music and dance has garnered increasing research interest. While recent advancements have led to notable progress in this area, existing methods predominantly operate in an offline manner—that is, they require access to the entire dance sequence before generating corresponding camera motions. This constraint renders them impractical for real-time applications, particularly in live stage performances, where immediate responsiveness is essential. To address this limitation, we introduce a more practical yet challenging task: online camera movement synthesis, in which camera trajectories must be generated using only the current and preceding segments of dance and music. In this paper, we propose TemMEGA (Temporal Masked Generative Modeling), a unified framework capable of handling both online and offline camera movement generation. TemMEGA consists of three key components. First, a discrete camera tokenizer encodes camera motions as discrete tokens via a discrete quantization scheme. Second, a consecutive memory encoder captures historical context by jointly modeling long-and short-term temporal dependencies across dance and music sequences. Finally, a temporal conditional masked transformer is employed to predict future camera motions by leveraging masked token prediction. Extensive experimental evaluations demonstrate the effectiveness of our TemMEGA, highlighting its superiority in both online and offline camera movement synthesis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li 等NeurIPS 2021 · 被引用 196 次
- Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic MemoryLi Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin 等CVPR 2022 · 被引用 170 次
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 被引用 151 次
相关 Paper
- DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera SynthesisZixuan Wang, Jiayi Li, Xiaoyu Qin, Shikun Sun 等ACM MM 2024 · 被引用 5 次
- DanceCamera3D: 3D Camera Movement Synthesis with Music and DanceZixuan Wang, Jia Jia, Shikun Sun, Haozhe Wu 等CVPR 2024 · 被引用 5 次
- Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion PriorForam Niravbhai Shah, Parshwa Shah, Muhammad Usama Saleem, Ekkasit Pinyoanuntapong 等AAAI 2026 · 被引用 4 次
- DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked ModelingAnindita Ghosh, Bing Zhou, Rishabh Dabral, Jian Wang 等SIGGRAPH 2025 · 被引用 11 次
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video GenerationKaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng 等SIGGRAPH 2026 · 被引用 3 次
