Real-time Object Detection for Streaming Perception
Jinrong Yang, Songtao Liu, Zeming Li, Xiaoping Li, Jian Sun
摘要
Autonomous driving requires the model to perceive the environment and (re)act within a low latency for safety. While past works ignore the inevitable changes in the environment after processing, streaming perception is proposed to jointly evaluate the latency and accuracy into a single metric for video online perception. In this paper, instead of searching trade-offs between accuracy and speed like previous works, we point out that endowing real-time models with the ability to predict the future is the key to dealing with this problem. We build a simple and effective framework for streaming perception. It equips a novel Dual-Flow Perception module (DFP), which includes dynamic and static flows to capture the moving trend and basic detection feature for streaming prediction. Further, we introduce a Trend-Aware Loss (TAL) combined with a trend factor to generate adaptive weights for objects with different moving speeds. Our simple method achieves competitive performance on Argoverse-HD dataset and improves the AP by 4.9% compared to the strong baseline, validating its effectiveness. Our code will be made available at https: //github.com/yancie-yjr/StreamYOLO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- PVT++: A Simple End-to-End Latency-Aware Visual Tracking FrameworkBowen Li, Ziyuan Huang, Junjie Ye, Yiming Li 等ICCV 2023 · 被引用 15 次
- Chanakya: Learning Runtime Decisions for Adaptive Real-Time PerceptionAnurag Ghosh, Vaibhav Balloli, Akshay Nambi, Aditya Singh 等NeurIPS 2023 · 被引用 10 次
- Real-time Stereo-based 3D Object Detection for Streaming PerceptionChangcai Li, Zonghua Gu, Gang Chen, Libo Huang 等NeurIPS 2024 · 被引用 3 次
- Transtreaming: Adaptive Delay-aware Transformer for Real-time Streaming PerceptionXiang Zhang, Yufei Cui, Chenchen Fu, Zihao Wang 等AAAI 2025 · 被引用 2 次
- Track-On: Transformer-based Online Point Tracking with MemoryGörkay Aydemir, Xiongyi Cai, Weidi Xie, Fatma GüneyICLR 2025
它引用的顶会 Paper13
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 被引用 1,030 次
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou 等ICCV 2019 · 被引用 211 次
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 被引用 45 次
- Predictive Feature Learning for Future Segmentation PredictionZihang Lin, Jiangxin Sun, Jianfang Hu, Qi-Zhi Yu 等ICCV 2021 · 被引用 18 次
相关 Paper
- SHARP: Short-Window Streaming for Accurate and Robust Prediction in Motion ForecastingAlexander Prutsch, Christian Fruhwirth-Reisinger, David Schinagl, Horst PosseggerCVPR 2026
- Are We Ready for Vision-Centric Driving Streaming Perception? The ASAP BenchmarkXiaofeng Wang, Zheng Zhu, Yunpeng Zhang, Guan Huang 等CVPR 2023
- 4DSegStreamer: Streaming 4D Panoptic Segmentation via Dual ThreadsLing Liu, Jun Tian, Li YiICCV 2025
- StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential EquationYining Shi, Kun Jiang, Ke Wang, Jiusi Li 等CVPR 2024
- ForeSight: Multi-View Streaming Joint Object Detection and Trajectory ForecastingSandro Papais, Letian Wang, Brian Cheong, Steven L. WaslanderICCV 2025 · 被引用 1 次
