EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User Interactions
Enhao Zhang, Maureen Daum, Dong He, Brandon Haynes, Ranjay Krishna, Magdalena Balazinska
Abstract
We introduce EQUI-VOCAL: a new system that automatically synthesizes queries over videos from limited user interactions. The user only provides a handful of positive and negative examples of what they are looking for. EQUI-VOCAL utilizes these initial examples and additional ones collected through active learning to efficiently synthesize complex user queries. Our approach enables users to find events without database expertise, with limited labeling effort, and without declarative specifications or sketches. Core to EQUI-VOCAL's design is the use of spatio-temporal scene graphs in its data model and query language and a novel query synthesis approach that works on large and noisy video data. Our system outperforms two baseline systems---in terms of F1 score, synthesis time, and robustness to noise---and can flexibly synthesize complex queries that the baselines do not support.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e87a5ac-a69f-4e69-ae34-60a334d9bb04Cited by top-tier papers5
- Agile Modeling: From Concept to Classifier in MinutesOtilia Stretcu, Edward Vendrow, Kenji Hata, Krishnamurthy Viswanathan et al.ICCV 2023 · 19 citations
- SketchQL: Video Moment Querying with a Visual Query InterfaceRenzhi Wu, Pramod Chunduri, Ali Payani, Xu Chu et al.SIGMOD 2025 · 6 citations
- Aero: Adaptive Query Processing of ML QueriesGaurav Tarlok Kakkar, Jiashen Cao, Aubhro Sengupta, Joy Arulraj et al.SIGMOD 2025 · 2 citations
- LOVO: Efficient Complex Object Query in Large-Scale Video DatasetsYuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu et al.ICDE 2025 · 2 citations
- Craw: A Unified and Efficient Querying Framework for Large-Scale Video DatasetsZiqi Zhou, Hanjian Jiang, Zihao Zeng, Xupuzhe Shao et al.VLDB 2026
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 103 citations
- Learning Optimal Tree Models under Beam SearchJingwei Zhuo, Ziru Xu, Wei Dai, Han Zhu et al.ICML 2020 · 72 citations
Related papers
- Self-Enhancing Video Data Management System for Compositional Events with Large Language ModelsEnhao Zhang, Nicole Sullivan, Brandon Haynes, Ranjay Krishna et al.SIGMOD 2025 · 4 citations
- Synthesizing Trajectory Queries from ExamplesStephen Mell, Favyen Bastani, Steve Zdancewic, Osbert BastaniCAV 2023 · 5 citations
- VOCALExplore: Pay-as-You-Go Video Data Exploration and Model BuildingMaureen Daum, Enhao Zhang, Dong He, Stephen Mussmann et al.VLDB 2023 · 7 citations
- IntentVizor: Towards Generic Query Guided Interactive Video SummarizationGuande Wu, Jianzhe Lin, Cláudio T. SilvaCVPR 2022 · 36 citations
- Synthesizing Document Database Queries Using Collection AbstractionsQikang Liu, Yang He, Yanwen Cai, Byeongguk Kwak et al.ICSE 2025
