SAND: A New Programming Abstraction for Video-based Deep Learning
Juncheol Ye, Seungkook Lee, Hwijoon Lim, Jihyuk Lee, Uitaek Hong, Youngjin Kwon, Dongsu Han
摘要
Video-based deep learning (VDL) is increasingly used across diverse applications and has become highly popular, but it faces significant challenges in preprocessing highly compressed video data. Preprocessing pipelines are complex, requiring extensive engineering effort, and introduce computational bottlenecks, with latency exceeding GPU training time. Existing solutions partially mitigate these issues but remain inefficient and resource-constrained.
We present SAND, a framework for VDL that integrates system-level optimizations to simplify the preprocessing pipeline and maximize resource efficiency. First, SAND introduces a view abstraction that encapsulates key preprocessing stages into virtualized objects, eliminating the need for users to manage individual objects. Second, SAND maximizes reuse opportunities through efficient system-level object management, reducing the preprocessing overhead and improving GPU utilization. Evaluation across multiple VDL applications and diverse environments, including Raybased hyperparameter search and distributed data parallel training, shows GPU utilization improvements of up to 12.3× and 2.9× over CPU and GPU baselines, respectively, while reducing preprocessing code complexity from hundreds or thousands of lines to fewer than 10.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mixtera: A Data Plane for Foundation Model TrainingMaximilian Böther, Xiaozhe Yao, Tolga Kerimoglu, Dan Graur 等SIGMOD 2026 · 被引用 5 次
- Neural-Centric Video Processing Pipeline for Unified Multi-Task InferenceSeyeon Lee, Juncheol Ye, Jaehong Kim, Dongsu HanCVPR 2026
它引用的顶会 Paper16
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 被引用 522 次
相关 Paper
- EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized ViewsZhuangdi Xu, Gaurav Tarlok Kakkar, Joy Arulraj, Umakishore RamachandranSIGMOD 2022 · 被引用 26 次
- OTIF: Efficient Tracker Pre-processing over Large Video DatasetsFavyen Bastani, Samuel MaddenSIGMOD 2022 · 被引用 20 次
- Video Decomposition Prior: Editing Videos Layer by LayerGaurav Shrivastava, Ser-Nam Lim, Abhinav ShrivastavaICLR 2024 · 被引用 11 次
- Accelerating Streaming Video Large Language Models via Hierarchical Token CompressionYiyu Wang, Xuyang Liu, Xiyan Gui, Xinying Lin 等CVPR 2026 · 被引用 40 次
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation ReuseJinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim 等VLDB 2025 · 被引用 2 次
