Eventfulness for Interactive Video Alignment
Jiatian Sun, Longxiulin Deng, Triantafyllos Afouras, Andrew Owens, Abe Davis
Abstract
Humans are remarkably sensitive to the alignment of visual events with other stimuli, which makes synchronization one of the hardest tasks in video editing. A key observation of our work is that most of the alignment we do involves salient localizable events that occur sparsely in time. By learning how to recognize these events, we can greatly reduce the space of possible synchronizations that an editor or algorithm has to consider. Furthermore, by learning descriptors of these events that capture additional properties of visible motion, we can build active tools that adapt their notion of eventfulness to a given task as they are being used. Rather than learning an automatic solution to one specific problem, our goal is to make a much broader class of interactive alignment tasks significantly easier and less time-consuming. We show that a suitable visual event descriptor can be learned entirely from stochastically-generated synthetic video. We then demonstrate the usefulness of learned and adaptive eventfulness by integrating it in novel interactive tools for applications including audio-driven time warping of video and the extraction and application of sound effects across different videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 22 citations
- EditDuet: A Multi-Agent System for Video Non-Linear EditingMarcelo Sandoval-Castañeda, Bryan C. Russell, Josef Sivic, Gregory Shakhnarovich et al.SIGGRAPH 2025 · 6 citations
- Video-Guided Foley Sound Generation with Multimodal ControlsZiyang Chen, Prem Seetharaman, Bryan C. Russell, Oriol Nieto et al.CVPR 2025
- Supervising Sound Localization by In-the-wild EgomotionAnna Min, Ziyang Chen, Hang Zhao, Andrew OwensCVPR 2025
Builds on3
- ChoreoMaster: choreography-oriented music-driven dance synthesisKang Chen, Zhipeng Tan, Jin Lei, Song-Hai Zhang et al.SIGGRAPH 2021 · 73 citations
- SpeedNet: Learning the Speediness in VideosSagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri et al.CVPR 2020
- AutoFlow: Learning a Better Training Set for Optical FlowDeqing Sun, Daniel Vlasic, Charles Herrmann, Varun Jampani et al.CVPR 2021
Related papers
- Soundify: Matching Sound Effects to VideoDavid Chuan-En Lin, Anastasis Germanidis, Cristóbal Valenzuela, Yining Shi et al.UIST 2023 · 15 citations
- How to Learn a Domain-Adaptive Event Simulator?Daxin Gu, Jia Li, Yu Zhang, Yonghong TianACM MM 2021 · 8 citations
- MoSound: An Interactive Tool for Generative Sound Design in Motion GraphicsJialin Huang, Prem Seetharaman, Timothy Richard Langlois, Li-Yi Wei et al.CHI 2026 · 2 citations
- D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX MatchingJingyu Liu, Minquan Wang, Ye Ma, Bo Wang et al.AAAI 2025 · 4 citations
- Video to Events: Recycling Video Datasets for Event CamerasDaniel Gehrig, Mathias Gehrig, Javier Hidalgo-Carrió, Davide ScaramuzzaCVPR 2020
