EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User Interactions
Enhao Zhang, Maureen Daum, Dong He, Brandon Haynes, Ranjay Krishna, Magdalena Balazinska
摘要
We introduce EQUI-VOCAL: a new system that automatically synthesizes queries over videos from limited user interactions. The user only provides a handful of positive and negative examples of what they are looking for. EQUI-VOCAL utilizes these initial examples and additional ones collected through active learning to efficiently synthesize complex user queries. Our approach enables users to find events without database expertise, with limited labeling effort, and without declarative specifications or sketches. Core to EQUI-VOCAL's design is the use of spatio-temporal scene graphs in its data model and query language and a novel query synthesis approach that works on large and noisy video data. Our system outperforms two baseline systems---in terms of F1 score, synthesis time, and robustness to noise---and can flexibly synthesize complex queries that the baselines do not support.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Agile Modeling: From Concept to Classifier in MinutesOtilia Stretcu, Edward Vendrow, Kenji Hata, Krishnamurthy Viswanathan 等ICCV 2023 · 被引用 19 次
- SketchQL: Video Moment Querying with a Visual Query InterfaceRenzhi Wu, Pramod Chunduri, Ali Payani, Xu Chu 等SIGMOD 2025 · 被引用 6 次
- Aero: Adaptive Query Processing of ML QueriesGaurav Tarlok Kakkar, Jiashen Cao, Aubhro Sengupta, Joy Arulraj 等SIGMOD 2025 · 被引用 2 次
- LOVO: Efficient Complex Object Query in Large-Scale Video DatasetsYuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu 等ICDE 2025 · 被引用 2 次
- Craw: A Unified and Efficient Querying Framework for Large-Scale Video DatasetsZiqi Zhou, Hanjian Jiang, Zihao Zeng, Xupuzhe Shao 等VLDB 2026
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
- Learning Optimal Tree Models under Beam SearchJingwei Zhuo, Ziru Xu, Wei Dai, Han Zhu 等ICML 2020 · 被引用 72 次
相关 Paper
- Self-Enhancing Video Data Management System for Compositional Events with Large Language ModelsEnhao Zhang, Nicole Sullivan, Brandon Haynes, Ranjay Krishna 等SIGMOD 2025 · 被引用 4 次
- Synthesizing Trajectory Queries from ExamplesStephen Mell, Favyen Bastani, Steve Zdancewic, Osbert BastaniCAV 2023 · 被引用 5 次
- VOCALExplore: Pay-as-You-Go Video Data Exploration and Model BuildingMaureen Daum, Enhao Zhang, Dong He, Stephen Mussmann 等VLDB 2023 · 被引用 7 次
- IntentVizor: Towards Generic Query Guided Interactive Video SummarizationGuande Wu, Jianzhe Lin, Cláudio T. SilvaCVPR 2022 · 被引用 36 次
- Synthesizing Document Database Queries Using Collection AbstractionsQikang Liu, Yang He, Yanwen Cai, Byeongguk Kwak 等ICSE 2025
