VidSTR: Automatic Spatiotemporal Retargeting of Speech-Driven Video Compositions
Joshua Kong Yang, Mackenzie Leake, Jeff Huang, Stephen DiVerdi
摘要
Video editors often record multiple versions of a performance with minor differences. When they add graphics atop one video, they may wish to transfer those assets to another recording, but differences in performance, wordings, and timings can cause assets to no longer be aligned with the video content. Fixing this is a time-consuming, manual task. We present a technique which preserves the temporal and spatial alignment of the original composition when automatically retargeting speech-driven video compositions. It can transfer graphics between both similar and dissimilar performances, including those varying in speech and gesture. We use a large language model for transcript-based temporal alignment and integer programming for spatial alignment. Results from retargeting between 51 pairs of performances show that we achieve a temporal alignment success rate of 90% compared to hand-generated ground truth compositions. We demonstrate challenging scenarios, retargeting video compositions across different people, aspect ratios, and languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- GRIDS: Interactive Layout Design with Integer ProgrammingNiraj Ramesh Dayama, Kashyap Todi, Taru Saarelainen, Antti OulasvirtaCHI 2020 · 被引用 65 次
- Scout: Rapid Exploration of Interface Layout Alternatives through High-Level Design ConstraintsAmanda Swearngin, Chenglong Wang, Alannah Oleson, James Fogarty 等CHI 2020 · 被引用 58 次
- Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup VideosAnh Truong, Peggy Chi, David Salesin, Irfan Essa 等CHI 2021 · 被引用 57 次
- ReelFramer: Human-AI Co-Creation for News-to-Video TranslationSitong Wang, Samia Menon, Tao Long, Keren Henderson 等CHI 2024 · 被引用 47 次
- Visual Captions: Augmenting Verbal Communication with On-the-fly VisualsXingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal 等CHI 2023 · 被引用 45 次
相关 Paper
- Audio-driven Neural Gesture Reenactment with Video Motion GraphsYang Zhou, Jimei Yang, Dingzeyu Li, Jun Saito 等CVPR 2022 · 被引用 22 次
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsSicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li 等ACM MM 2023 · 被引用 17 次
- Do You Have Freestyle? Expressive Humanoid Locomotion via Audio ControlZhe Li, Cheng Chi, Yangyang Wei, Boan Zhu 等CVPR 2026 · 被引用 13 次
- Semantics-Aware Motion Retargeting with Vision-Language ModelsHaodong Zhang, Zhike Chen, Haocheng Xu, Lei Hao 等CVPR 2024 · 被引用 8 次
- MoVer: Motion Verification for Motion Graphics AnimationsJiaju Ma, Maneesh AgrawalaSIGGRAPH 2025 · 被引用 8 次
