AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference
Hangfeng Liang, Yutao Hu, Yanhan Hu, Xiaohan Wu, WENQI SHAO, Ying Fu
Abstract
Low-light video enhancement (LLVE) remains a challenging task due to severe information degradation under low-illumination conditions. Recent multimodal approaches have significantly improved enhancement performance by incorporating auxiliary modalities, such as event streams and infrared images. However, these methods typically assume the availability of these modalities at inference, which is often not feasible in real-world scenarios. To solve this problem, in this work, we propose AMNet, a unified multimodal framework for LLVE, to support flexible modality-agnostic inference, where auxiliary modalities may be unavailable. To address the issue of modality absence, we introduce a Spatial-Spectral Dual-Gated Translator that learns the correspondence between auxiliary modalities and RGB inputs, producing implicit auxiliary representations to support the robust enhancement. Additionally, to fully facilitate the learning of cross-modal correspondence, we conduct large-scale multimodal pretraining based on the RGB-only dataset with synthetic auxiliary modalities. Extensive experiments demonstrate that AMNet could handle arbitrary inference-time modality combinations and exhibits superior performance for LLVE under modality absence conditions. Code and models are available on the project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31aaf1e8-2cb2-4a1a-9abb-69df2d3b89c1Builds on22
- Retinexformer: One-stage Retinex-based Transformer for Low-light Image EnhancementYuanhao Cai, Hao Bian, Jing Lin, Haoqian Wang et al.ICCV 2023 · 615 citations
- Seeing Motion in the DarkChen Chen, Qifeng Chen, Minh N. Do, Vladlen KoltunICCV 2019 · 315 citations
- Seeing Dynamic Scene in the Dark: A High-Quality Video Dataset with Mechatronic AlignmentRuixing Wang, Xiaogang Xu, Chi-Wing Fu, Jiangbo Lu et al.ICCV 2021 · 160 citations
- Unidentified Video Objects: A Benchmark for Dense, Open-World SegmentationWeiyao Wang, Matt Feiszli, Heng Wang, Du TranICCV 2021 · 151 citations
- Robust Monocular Depth Estimation under Challenging ConditionsStefano Gasperini, Nils Morbitzer, HyunJun Jung, Nassir Navab et al.ICCV 2023 · 87 citations
Related papers
- Event-Guided Consistent Video Enhancement with Modality-Adaptive Diffusion PipelineKanghao Chen, Zixin Zhang, Guoqiang Liang, Lutao Jiang et al.NeurIPS 2025 · 2 citations
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 84 citations
- Coherent Event Guided Low-Light Video EnhancementJinxiu Liang, Yixin Yang, Boyu Li, Peiqi Duan et al.ICCV 2023 · 54 citations
- CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training FrameworkWentao Wu, Xiao Wang, Chenglong Li, Bo Jiang et al.ACM MM 2025 · 2 citations
- Low-Light Video Enhancement with Synthetic Event GuidanceLin Liu, Junfeng An, Jianzhuang Liu, Shanxin Yuan et al.AAAI 2023 · 51 citations
