Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
Qiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen, Ke Lin, Jiahao Wang, Boyuan Jiang, Haotian Yang, Mingwu Zheng, Xin Tao, Fei Yang, Pengfei Wan, Di Zhang
Abstract
Panda-70M Unfiltered "A woman is cracking eggs into a bowl of spinach in the kitchen." Panda-70M Koala-36M "A woman is standing in a modern kitchen, engaging in a conversation or explaining something while gesturing with her hands. She is wearing a black top and has long, braided hair. The kitchen is well-lit with warm lighting, and there are various kitchen items on the counter, including a vase with red flowers, a bowl of eggs, and a green bowl. The woman appears to be in a cheerful mood, smiling……" Koala-36M "a person preparing a dish in a kitchen setting. The person is seen cracking eggs into a bowl filled with spinach leaves. The scene is focused on the hands and the bowl, with various kitchen items like a bottle of oil, a container of salt, and a teapot visible in the background. The person's hands are the main focus, showing the careful and deliberate action of cracking the eggs and adding them to the ……" "A shirtless man flexing his muscles in front of a crowd." Koala-36M "A muscular individual with a tattooed torso and arms is standing in front of a microphone, holding a piece of paper. The person is wearing a white tank top and appears to be in a celebratory or victorious mood, as indicated by their raised fists and the expression of triumph on their face. The background suggests a sports event, specifically a boxing match, as indicated by the presence of a microphone……" Koala-36M "Two men are engaged in a handshake, with one of them flexing his muscles. The man on the left has a heavily tattooed arm, with visible ink on his forearm and bicep. He is wearing a black sleeveless shirt and has a yellow wristband. The man on the right has a muscular build, with a tattoo on his right arm and a cap on his head. He is shirtless, revealing his well-defined muscles ……" VTSS 3.72 VTSS 4.13 VTSS 3.69 VTSS 3.56 Panda-70M Unfiltered Koala-36M Filter out (freeze-frame video) VTSS 1.05 VTSS 1.69 Koala-36M Filter out (overexposed video) Panda-70M Figure 1. Comparison between Koala-36M and Panda-70M. We propose a large-scale, high-quality dataset that significantly enhances the consistency between multiple conditions and video content. Koala-36M features more accurate temporal splitting, more detailed captions, and improved video filtering based on the proposed Video Training Suitability Score (VTSS).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 416747b7-ee2d-4736-ba95-71a8184fcb15Cited by top-tier papers49
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park et al.CVPR 2026 · 144 citations
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng et al.NeurIPS 2025 · 98 citations
- WISA: World simulator assistant for physics-aware text-to-video generationJing Wang, Ao Ma, Ke Cao, Jun Zheng et al.NeurIPS 2025 · 93 citations
- Let Them Talk: Audio-Driven Multi-Person Conversational Video GenerationZhe Kong, Feng Gao, Yong Zhang, Zhuoliang Kang et al.NeurIPS 2025 · 73 citations
- SpatialVID: A Large-Scale Video Dataset with Spatial AnnotationsJiahao Wang, Yufeng Yuan, Rujie Zheng, Youtian Lin et al.CVPR 2026 · 72 citations
Builds on19
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel et al.ICCV 2023 · 800 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
Related papers
- Panda-70M: Captioning 70M Videos with Multiple Cross-Modality TeachersTsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka et al.CVPR 2024
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosChiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha et al.CVPR 2025
- HD-EPIC: A Highly-Detailed Egocentric Video DatasetToby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara et al.CVPR 2025
- HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language ModelsXiao Wang, Jingyun Hua, Weihong Lin, Yuanxing Zhang et al.ACL 2025 · 1 citation
- Spatiotemporal Skip Guidance for Enhanced Video Diffusion SamplingJunha Hyung, Kinam Kim, Susung Hong, Min-Jung Kim et al.CVPR 2025
