Zeus: Efficiently Localizing Actions in Videos using Reinforcement Learning
Pramod Chunduri, Jaeho Bang, Yao Lu, Joy Arulraj
摘要
Detection and localization of actions in videos is an important problem in practice. State-of-the-art video analytics systems are unable to efficiently and effectively answer such action queries because actions often involve a complex interaction between objects and are spread across a sequence of frames; detecting and localizing them requires computationally expensive deep neural networks. It is also important to consider the entire sequence of frames to answer the query effectively.
In this paper, we present Zeus, a video analytics system tailored for answering action queries. We present a novel technique for efficiently answering these queries using deep reinforcement learning. Zeus trains a reinforcement learning agent that learns to adaptively modify the input video segments that are subsequently sent to an action classification network. The agent alters the input segments along three dimensions -sampling rate, segment length, and resolution. To meet the user-specified accuracy target, Zeus's query optimizer trains the agent based on an accuracy-aware, aggregate reward function. Evaluation on three diverse video datasets shows that Zeus outperforms state-of-the-art frame-and window-based filtering techniques by up to 22.1× and 4.7×, respectively. It also consistently meets the user-specified accuracy target across all queries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SEIDEN: Revisiting Query Processing in Video Database SystemsJaeho Bang, Gaurav Tarlok Kakkar, Pramod Chunduri, Subrata Mitra 等VLDB 2023 · 被引用 24 次
- Extract-Transform-Load for Video StreamsFerdinand Kossmann, Ziniu Wu, Eugenie Lai, Nesime Tatbul 等VLDB 2023 · 被引用 21 次
- EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User InteractionsEnhao Zhang, Maureen Daum, Dong He, Brandon Haynes 等VLDB 2023 · 被引用 18 次
- SketchQL: Video Moment Querying with a Visual Query InterfaceRenzhi Wu, Pramod Chunduri, Ali Payani, Xu Chu 等SIGMOD 2025 · 被引用 6 次
- Predictive and Near-Optimal Sampling for View Materialization in Video DatabasesYanchao Xu, Dongxiang Zhang, Shuhao Zhang, Sai Wu 等SIGMOD 2024 · 被引用 5 次
它引用的顶会 Paper5
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
- MIRIS: Fast Object Track Queries in VideoFavyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan 等SIGMOD 2020 · 被引用 68 次
- GOGGLES: Automatic Image Labeling with Affinity CodingNilaksh Das, Sanya Chaba, Renzhi Wu, Sakshi Gandhi 等SIGMOD 2020 · 被引用 22 次
- ODIN: Automated Drift Detection and Recovery in Video AnalyticsAbhijit Suprem, Joy Arulraj, Calton Pu, João Eduardo FerreiraVLDB 2020
- Straight to the Point: Fast-Forwarding Videos via Reinforcement Learning Using Textual DataWashington L. S. Ramos, Michel Melo Silva, Edson R. Araujo, Leandro Soriano Marcolino 等CVPR 2020
相关 Paper
- Batch Adaptative Streaming for Video AnalyticsLei Zhang, Yuqing Zhang, Ximing Wu, Fangxin Wang 等INFOCOM 2022 · 被引用 24 次
- Top-K Deep Video Analytics: A Probabilistic ApproachZiliang Lai, Chenxia Han, Chris Liu, Pengfei Zhang 等SIGMOD 2021 · 被引用 7 次
- STRONG: Spatio-Temporal Reinforcement Learning for Cross-Modal Video Moment LocalizationDa Cao, Yawen Zeng, Meng Liu, Xiangnan He 等ACM MM 2020 · 被引用 47 次
- Fast Template Matching and Update for Video Object Tracking and SegmentationMingjie Sun, Jimin Xiao, Eng Gee Lim, Bingfeng Zhang 等CVPR 2020
- CASVA: Configuration-Adaptive Streaming for Live Video AnalyticsMiao Zhang, Fangxin Wang, Jiangchuan LiuINFOCOM 2022 · 被引用 69 次
