VidSTR: Automatic Spatiotemporal Retargeting of Speech-Driven Video Compositions
Joshua Kong Yang, Mackenzie Leake, Jeff Huang, Stephen DiVerdi
Abstract
Video editors often record multiple versions of a performance with minor differences. When they add graphics atop one video, they may wish to transfer those assets to another recording, but differences in performance, wordings, and timings can cause assets to no longer be aligned with the video content. Fixing this is a time-consuming, manual task. We present a technique which preserves the temporal and spatial alignment of the original composition when automatically retargeting speech-driven video compositions. It can transfer graphics between both similar and dissimilar performances, including those varying in speech and gesture. We use a large language model for transcript-based temporal alignment and integer programming for spatial alignment. Results from retargeting between 51 pairs of performances show that we achieve a temporal alignment success rate of 90% compared to hand-generated ground truth compositions. We demonstrate challenging scenarios, retargeting video compositions across different people, aspect ratios, and languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb9b7b42-77ea-4540-be86-bd76bb48bb87Builds on13
- GRIDS: Interactive Layout Design with Integer ProgrammingNiraj Ramesh Dayama, Kashyap Todi, Taru Saarelainen, Antti OulasvirtaCHI 2020 · 65 citations
- Scout: Rapid Exploration of Interface Layout Alternatives through High-Level Design ConstraintsAmanda Swearngin, Chenglong Wang, Alannah Oleson, James Fogarty et al.CHI 2020 · 58 citations
- Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup VideosAnh Truong, Peggy Chi, David Salesin, Irfan Essa et al.CHI 2021 · 57 citations
- ReelFramer: Human-AI Co-Creation for News-to-Video TranslationSitong Wang, Samia Menon, Tao Long, Keren Henderson et al.CHI 2024 · 47 citations
- Visual Captions: Augmenting Verbal Communication with On-the-fly VisualsXingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal et al.CHI 2023 · 45 citations
Related papers
- Audio-driven Neural Gesture Reenactment with Video Motion GraphsYang Zhou, Jimei Yang, Dingzeyu Li, Jun Saito et al.CVPR 2022 · 22 citations
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsSicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li et al.ACM MM 2023 · 17 citations
- Do You Have Freestyle? Expressive Humanoid Locomotion via Audio ControlZhe Li, Cheng Chi, Yangyang Wei, Boan Zhu et al.CVPR 2026 · 13 citations
- Semantics-Aware Motion Retargeting with Vision-Language ModelsHaodong Zhang, Zhike Chen, Haocheng Xu, Lei Hao et al.CVPR 2024 · 8 citations
- MoVer: Motion Verification for Motion Graphics AnimationsJiaju Ma, Maneesh AgrawalaSIGGRAPH 2025 · 8 citations
