VOCALExplore: Pay-as-You-Go Video Data Exploration and Model Building
Maureen Daum, Enhao Zhang, Dong He, Stephen Mussmann, Brandon Haynes, Ranjay Krishna, Magdalena Balazinska
Abstract
We introduce VOCALExplore, a system designed to support users in building domain-specific models over video datasets. VOCALExplore supports interactive labeling sessions and trains models using user-supplied labels. VOCALExplore maximizes model quality by automatically deciding how to select samples based on observed skew in the collected labels. It also selects the optimal video representations to use when training models by casting feature selection as a rising bandit problem. Finally, VOCALExplore implements optimizations to achieve low latency without sacrificing model performance. We demonstrate that VOCALExplore achieves close to the best possible model quality given candidate acquisition functions and feature extractors, and it does so with low visible latency ( 1 second per iteration) and no expensive preprocessing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Agile Modeling: From Concept to Classifier in MinutesOtilia Stretcu, Edward Vendrow, Kenji Hata, Krishnamurthy Viswanathan et al.ICCV 2023 · 19 citations
- EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User InteractionsEnhao Zhang, Maureen Daum, Dong He, Brandon Haynes et al.VLDB 2023 · 18 citations
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation ReuseJinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim et al.VLDB 2025 · 2 citations
- LOVO: Efficient Complex Object Query in Large-Scale Video DatasetsYuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu et al.ICDE 2025 · 2 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 103 citations
- MIRIS: Fast Object Track Queries in VideoFavyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan et al.SIGMOD 2020 · 68 citations
Related papers
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- Understanding Human Preferences: Towards More Personalized Video to Text GenerationYihan Wu, Ruihua Song, Xu Chen, Hao Jiang et al.WWW 2024 · 6 citations
- Frame-Voyager: Learning to Query Frames for Video Large Language ModelsSicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen et al.ICLR 2025
- CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLMSon Nguyen, Xinyuan Liu, Ransalu SenanayakeICML 2026
- Self-Enhancing Video Data Management System for Compositional Events with Large Language ModelsEnhao Zhang, Nicole Sullivan, Brandon Haynes, Ranjay Krishna et al.SIGMOD 2025 · 4 citations
