What Are You Doing? A Closer Look at Controllable Human Video Generation
Emanuele Bugliarello, Anurag Arnab, Roni Paiss, Christy Koh, Pieter-Jan Kindermans, Cordelia Schmid
Abstract
High-quality benchmarks are crucial for driving progress in machine learning research. However, despite the growing interest in video generation, there is no comprehensive dataset to evaluate human synthesis. Humans can perform a wide variety of actions and interactions, but existing datasets, like TikTok, TED-Talks, and HumanVid, lack the diversity and complexity to fully capture the capabilities of video generation models. We close this gap by introducing 'What Are You Doing?' (WYD): a new benchmark for fine-grained evaluation of controllable image-to-video generation of humans. WYD consists of 1,544 captioned videos that have been meticulously collected and annotated with fine-grained categories. These allow us to systematically measure performance across 9 aspects of human generation, including actions, interactions and motion. We also propose an evaluation framework, where we adapt existing metrics for better human- and video-level assessment, as shown by human preference. Equipped with our dataset and metrics, we perform in-depth analyses of state-of-the-art open-source models in controllable image-to-video generation, showing how WYD provides novel insights about their capabilities. We release data and code to drive progress in human video generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4332c714-f03e-4309-9d18-a67dada1874eCited by top-tier papers2
- DanceTogether: Generating Interactive Multi-Person Video without Identity DriftingJunhao Chen, Mingjin Chen, Jianjin Xu, Xiang Li et al.ICLR 2026
- Towards Fine-Grained Human Motion Video CaptioningGuorui Song, Guocun Wang, Zhe Huang, Jing Lin et al.ACM MM 2025
Builds on46
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
Related papers
- IF-VidCap: Can Video Caption Models Follow Instructions?Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei et al.ICLR 2026 · 7 citations
- VBench: Comprehensive Benchmark Suite for Video Generative ModelsZiqi Huang, Yinan He, Jiashuo Yu, Fan Zhang et al.CVPR 2024
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- OSCBench: Benchmarking Object State Change in Text-to-Video GenerationXianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li et al.ACL 2026 · 2 citations
- EvalCrafter: Benchmarking and Evaluating Large Video Generation ModelsYaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang et al.CVPR 2024
