Instant Video Models: Universal Adapters for Stabilizing Image-Based Networks
Matthew Dutson, Nathan Labiosa, Yin Li, Mohit Gupta
Abstract
When applied sequentially to video, frame-based networks often exhibit temporal inconsistency-for example, outputs that flicker between frames. This problem is amplified when the network inputs contain time-varying corruptions. In this work, we introduce a general approach for adapting frame-based models for stable and robust inference on video. We describe a class of stability adapters that can be inserted into virtually any architecture and a resource-efficient training process that can be performed with a frozen base network. We introduce a unified conceptual framework for describing temporal stability and corruption robustness, centered on a proposed accuracy-stability-robustness loss. By analyzing the theoretical properties of this loss, we identify the conditions where it produces well-behaved stabilizer training. Our experiments validate our approach on several vision tasks including denoising (NAFNet), image enhancement (HDRNet), monocular depth (Depth Anything v2), and semantic segmentation (DeepLabv3+). Our method improves temporal stability and robustness against a range of image corruptions (including compression artifacts, noise, and adverse weather), while preserving or improving the quality of predictions.
Several prior works have proposed video-centric models with improved temporal consistency [5,44,24,74,60,72]. However, these methods are often narrowly designed for one or a few tasks and require costly training on large-scale video datasets. Consequently, they lack the flexibility to leverage the extensive ecosystem of frame-based imaging and perception models. Further, few explicitly address robustness to transient corruptions or other challenging conditions.
In this work, we improve the temporal consistency and robustness of pre-trained, image-based models across various tasks. One of our primary challenges is that increasing stability may result in over-smoothing, which can, in turn, reduce accuracy. We conceptually explore the tradeoffs between output quality, corruption robustness, and temporal consistency, and introduce a unified accuracy-robustness-stability loss to balance these objectives. We provide a theoretical analysis of this loss and identify strategies to avoid "over-smoothing reality," such that there is no incentive for predictions to be smoother than the true scene dynamics.
Guided by this analysis, we propose a class of versatile stabilization adapters (Figure 1 middle). These adapters generate control signals, based on recent spatiotemporal context, that modulate changes to the model's features and output. By operating in both the feature and output spaces, we allow the adapters to model stability wherever it exists in the visual hierarchy. This property is important for high-level vision tasks, where stability is often best described in a feature space.
Our method offers several key benefits. First, our stabilization adapters are lightweight and modular, and do not require modifying the original model parameters. Second, our adapters operate causally; stabilized outputs depend only on current and past inputs-a feature that is critical for processing streaming video in latency-sensitive applications. Third, our approach is compatible with both lowlevel tasks, where stability can be described in terms of pixel values, and higher-level tasks, where stability occurs at the level of scene semantics. Finally, our method naturally enhances robustness to transient corruptions, without requiring explicit corruption modeling.
We evaluate our method on a range of tasks: denoising, image enhancement, monocular depth estimation, and semantic segmentation (Figure 1 middle-right). We also demonstrate improved robustness against various transient corruptions, including noise, dropped patches, elastic deformations, compression artifacts, and adverse weather (Figure 1 bottom). In most cases, these improvements do not reduce accuracy-on the contrary, we often see significant improvements in task metrics. Overall, our experiments establish the flexibility and practicality of our approach.
2 Related Work Corruption robustness. Several prior works have addressed robustness against input corruptions. Hendrycks and Dietterich [23] propose metrics for measuring the robustness of image classifiers against common corruptions (e.g., compression artifacts or weather); their metrics inspire our definitions in Section 3. In general, natural corruptions have received less attention [14] from the vision community than adversarial corruptions [1,19,46,48,59,61,69,71], although a handful of methods and benchmarks exist [3,34,41,47,70]. Like these works, our paper emphasizes robustness to naturally occurring corruptions rather than worst-case adversarial perturbations.
Frame-to-frame flickering is a significant problem for low-level image enhancement models; as such, there have been several works that improve temporal consistency for these tasks [5,33,37,77,80]. Blind video temporal consistency m
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5165d931-6d09-4342-a7eb-c44c8f35f240Builds on18
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Arbitrary Video Style Transfer via Multi-Channel CorrelationYingying Deng, Fan Tang, Weiming Dong, Haibin Huang et al.AAAI 2021 · 197 citations
- Rethinking Softmax Cross-Entropy Loss for Adversarial RobustnessTianyu Pang, Kun Xu, Yinpeng Dong, Chao Du et al.ICLR 2020 · 176 citations
Related papers
- Exploring perceptual straightness in learned visual representationsAnne Harrington, Vasha DuTell, Ayush Tewari, Mark Hamilton et al.ICLR 2023
- Video Dynamics Prior: An Internal Learning Approach for Robust Video EnhancementsGaurav Shrivastava, Ser Nam Lim, Abhinav ShrivastavaNeurIPS 2023 · 14 citations
- Benchmarking the Robustness of Temporal Action Detection Models Against Temporal CorruptionsRunhao Zeng, Xiaoyong Chen, Jiaming Liang, Huisi Wu et al.CVPR 2024 · 6 citations
- Temporal Denoising Mask Synthesis Network for Learning Blind Video Temporal ConsistencyYifeng Zhou, Xing Xu, Fumin Shen, Lianli Gao et al.ACM MM 2020 · 9 citations
- VersVideo: Leveraging Enhanced Temporal Diffusion Models for Versatile Video GenerationJinxi Xiang, Ricong Huang, Jun Zhang, Guanbin Li et al.ICLR 2024 · 4 citations
