E²I-VRWKV: Explicit EPI-Representation and Interaction-Aware Vision-RWKV for Light Field Semantic Segmentation
Wei Zhang, Chen Jia, Xu Cheng, Fan Shi, Hui Liu, Shengyong Chen
Abstract
Pixel-level semantic segmentation of 4D light field (LF) data remains a considerable challenge, primarily due to the conflict between modeling complex spatial-angular dependencies and maintaining linear computational efficiency. Current linear models like VRWKV offer scalability but often fail to capture intrinsic geometric structures, leading to the structural collapse of Epipolar Plane Image (EPI) cues. To overcome these limitations, we propose E 2 I-VRWKV, an explicit EPI-Representation and Interaction-aware network that generates high-quality segmentation maps by embedding explicit geometric priors into a linear-complexity backbone. Specifically, we introduce the Light Field Epipolar-Aware Cross-Modal Attention (LF-ECMA) block.
The key innovation lies in the integration of an EPI Geometric Prior Generator, which explicitly extracts disparity-sensitive biases to enforce geometric consistency, and a Geometric-Context Gating (GC-Gate) mechanism. This mechanism functions as a geometrically modulated aperture to dynamically calibrate the fusion of spatial and angular manifolds. Experiments on the Ur-banLF benchmark demonstrate that our method outperforms other state-of-the-art (SOTA) methods, achieving 86.55% mIoU on UrbanLF-Real while maintaining an improved balance between accuracy and linear efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b0d6ac6-0a4a-4a12-aea5-07ff22762171Builds on10
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- RTFormer: Efficient Design for Real-Time Semantic Segmentation with TransformerJian Wang, Chenhui Gou, Qiman Wu, Haocheng Feng et al.NeurIPS 2022 · 207 citations
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu et al.ICML 2024 · 64 citations
- Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesYuchen Duan, Weiyun Wang, Zhe Chen, Xizhou Zhu et al.ICLR 2025 · 10 citations
Related papers
- Epipolar Consistency-based Network for Structure-Aware LF Semantic SegmentationChen Gao, Youfang Lin, Wenbin Wang, Shuo ZhangACM MM 2025
- Combining Implicit-Explicit View Correlation for Light Field Semantic SegmentationRuixuan Cong, Da Yang, Rongshan Chen, Sizhe Wang et al.CVPR 2023
- Learning Non-Local Spatial-Angular Correlation for Light Field Image Super-ResolutionZhengyu Liang, Yingqian Wang, Longguang Wang, Jungang Yang et al.ICCV 2023 · 72 citations
- Learning Spatial-angular Fusion for Compressive Light Field Imaging in a Cycle-consistent FrameworkXianqiang Lyu, Zhiyu Zhu, Mantang Guo, Jing Jin et al.ACM MM 2021 · 8 citations
- Epipolar Consistent Attention Aggregation Network for Unsupervised Light Field Disparity EstimationChen Gao, Shuo Zhang, Youfang LinICCV 2025 · 3 citations
