VOCALExplore: Pay-as-You-Go Video Data Exploration and Model Building
Maureen Daum, Enhao Zhang, Dong He, Stephen Mussmann, Brandon Haynes, Ranjay Krishna, Magdalena Balazinska
摘要
We introduce VOCALExplore, a system designed to support users in building domain-specific models over video datasets. VOCALExplore supports interactive labeling sessions and trains models using user-supplied labels. VOCALExplore maximizes model quality by automatically deciding how to select samples based on observed skew in the collected labels. It also selects the optimal video representations to use when training models by casting feature selection as a rising bandit problem. Finally, VOCALExplore implements optimizations to achieve low latency without sacrificing model performance. We demonstrate that VOCALExplore achieves close to the best possible model quality given candidate acquisition functions and feature extractors, and it does so with low visible latency ( 1 second per iteration) and no expensive preprocessing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Agile Modeling: From Concept to Classifier in MinutesOtilia Stretcu, Edward Vendrow, Kenji Hata, Krishnamurthy Viswanathan 等ICCV 2023 · 被引用 19 次
- EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User InteractionsEnhao Zhang, Maureen Daum, Dong He, Brandon Haynes 等VLDB 2023 · 被引用 18 次
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation ReuseJinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim 等VLDB 2025 · 被引用 2 次
- LOVO: Efficient Complex Object Query in Large-Scale Video DatasetsYuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu 等ICDE 2025 · 被引用 2 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 等EMNLP 2021 · 被引用 399 次
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
- MIRIS: Fast Object Track Queries in VideoFavyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan 等SIGMOD 2020 · 被引用 68 次
相关 Paper
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder 等ICML 2024 · 被引用 513 次
- Understanding Human Preferences: Towards More Personalized Video to Text GenerationYihan Wu, Ruihua Song, Xu Chen, Hao Jiang 等WWW 2024 · 被引用 6 次
- Frame-Voyager: Learning to Query Frames for Video Large Language ModelsSicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen 等ICLR 2025
- CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLMSon Nguyen, Xinyuan Liu, Ransalu SenanayakeICML 2026
- Self-Enhancing Video Data Management System for Compositional Events with Large Language ModelsEnhao Zhang, Nicole Sullivan, Brandon Haynes, Ranjay Krishna 等SIGMOD 2025 · 被引用 4 次
