Neural-Centric Video Processing Pipeline for Unified Multi-Task Inference
Seyeon Lee, Juncheol Ye, Jaehong Kim, Dongsu Han
Abstract
Videos are increasingly used as inputs to machine learning systems, where repeated decoding and processing across diverse downstream tasks dominate computational cost. However, existing video pipelines remain inefficient. Traditional codecs such as H.264 and H.265 are optimized for human perception and require full pixel decoding for every query, compressed-domain methods are tied to specific codec structures with limited flexibility, and machineoriented video coding approaches often rely on task-specific encoders and separate representations without supporting human visualization. We propose Neural Video Pipeline (NVP), a framework that leverages implicit neural representations to directly extract task-specific features from intermediate layers, eliminating pixel reconstruction overhead. NVP introduces lightweight micro adapters that map these features into the representation space of downstream models, bypassing both decoding and early-stage feature extraction. Across four representative tasks-image classification, object detection, action recognition, and segmentation-NVP reduces latency by up to 89.5% and inference FLOPs by up to 29.9%, while supporting multiple tasks using a single unified representation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb15e41e-7e9c-4720-b991-04b596389885Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- HiNeRV: Video Compression with Hierarchical Encoding-based Neural RepresentationHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2023 · 132 citations
Related papers
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren et al.NeurIPS 2021 · 430 citations
- HNeRV: A Hybrid Neural Representation for VideosHao Chen, Matthew Gwilliam, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2023
- Towards Scalable Neural Representation for Diverse VideosBo He, Xitong Yang, Hanyu Wang, Zuxuan Wu et al.CVPR 2023
- Neural Inter-Frame Compression for Video CodingAbdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher SchroersICCV 2019 · 207 citations
- AdaCLIP: Towards Pragmatic Multimodal Video RetrievalZhiming Hu, Angela Ning Ye, Salar Hosseini Khorasgani, Iqbal MohomedACM MM 2023 · 9 citations
