VideoDoodles: Hand-Drawn Animations on Videos with Scene-Aware Canvases
Emilie Yu, Kevin Blackburn-Matzen, Cuong Nguyen, Oliver Wang, Rubaiat Habib Kazi, Adrien Bousseau
Abstract
Fig. 1. Video doodles combine hand-drawn animations with video footage. Our interactive system eases the creation of this mixed media art by letting users place planar canvases in the scene which are then tracked in 3D. In this example, the inserted rainbow bridge exhibits correct perspective and occlusions, and the character's face and arms follow the tram as it runs towards the camera.
We present an interactive system to ease the creation of so-called video doodles -videos on which artists insert hand-drawn animations for entertainment or educational purposes. Video doodles are challenging to create because to be convincing, the inserted drawings must appear as if they were part of the captured scene. In particular, the drawings should undergo tracking, perspective deformations and occlusions as they move with respect to the camera and to other objects in the scene -visual effects that are difficult to reproduce with existing 2D video editing software. Our system supports these effects by relying on planar canvases that users position in a 3D scene reconstructed from the video. Furthermore, we present a custom tracking algorithm that allows users to anchor canvases to static or dynamic objects in the scene, such that the canvases move and rotate to follow the position and direction of these objects. When testing our system, novices could create a variety of short animated clips in a dozen of minutes, while professionals praised its speed and ease of use compared to existing tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3035e5e5-16c9-44c2-9cbc-7d3c79c191f1Cited by top-tier papers10
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo et al.NeurIPS 2023 · 555 citations
- RealityCanvas: Augmented Reality Sketching for Embedded and Responsive Scribble Animation EffectsZhijie Xia, Kyzyl Monteiro, Kevin Van, Ryo SuzukiUIST 2023 · 22 citations
- Tapnext: Tracking Any Point (Tap) as Next Token PredictionArtem Zholus, Carl Doersch, Yi Yang, Skanda Koppula et al.ICCV 2025 · 7 citations
- Multi-Object Sketch Animation by Scene Decomposition and Motion PlanningJingyu Liu, Zijie Xin, Yuhan Fu, Ruixiang Zhao et al.ICCV 2025 · 6 citations
- ViewCraft3D: High-fidelity and View-Consistent 3D Vector Graphics SynthesisChuang Wang, Haitao Zhou, Ling Luo, Qian YuNeurIPS 2025 · 5 citations
Builds on11
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented RealityRuofei Du, Eric Turner, Maksym Dzitsiuk, Luca Prasso et al.UIST 2020 · 145 citations
- RealitySketch: Embedding Responsive Graphics and Visualizations in AR through Dynamic SketchingRyo Suzuki, Rubaiat Habib Kazi, Li-Yi Wei, Stephen DiVerdi et al.UIST 2020 · 98 citations
- Pronto: Rapid Augmented Reality Video Prototyping Using Sketches and EnactionGermán Leiva, Cuong Nguyen, Rubaiat Habib Kazi, Paul AsenteCHI 2020 · 96 citations
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 69 citations
Related papers
- PoseTween: Pose-driven Tween AnimationJingyuan Liu, Hongbo Fu, Chiew-Lan TaiUIST 2020 · 19 citations
- SketchVideo: Sketch-based Video Generation and EditingFeng-Lin Liu, Hongbo Fu, Xintao Wang, Weicai Ye et al.CVPR 2025
- VideoCraft: A Mixed Reality-empowered Video Generation Workflow with Spatial Layer Editing for Concept Video CreationBoyu Li, Linping Yuan, Zeyu WangUIST 2025 · 3 citations
- VideoClipper: Rapid Prototyping with the "Editing-in-the-Camera" MethodWendy E. Mackay, Alexandre Battut, Germán Leiva, Michel Beaudouin-LafonCHI 2024 · 5 citations
- Looking Backward: Streaming Video-to-Video Translation with Feature BanksFeng Liang, Akio Kodaira, Chenfeng Xu, Masayoshi Tomizuka et al.ICLR 2025
