E²I-VRWKV: Explicit EPI-Representation and Interaction-Aware Vision-RWKV for Light Field Semantic Segmentation
Wei Zhang, Chen Jia, Xu Cheng, Fan Shi, Hui Liu, Shengyong Chen
摘要
Pixel-level semantic segmentation of 4D light field (LF) data remains a considerable challenge, primarily due to the conflict between modeling complex spatial-angular dependencies and maintaining linear computational efficiency. Current linear models like VRWKV offer scalability but often fail to capture intrinsic geometric structures, leading to the structural collapse of Epipolar Plane Image (EPI) cues. To overcome these limitations, we propose E 2 I-VRWKV, an explicit EPI-Representation and Interaction-aware network that generates high-quality segmentation maps by embedding explicit geometric priors into a linear-complexity backbone. Specifically, we introduce the Light Field Epipolar-Aware Cross-Modal Attention (LF-ECMA) block.
The key innovation lies in the integration of an EPI Geometric Prior Generator, which explicitly extracts disparity-sensitive biases to enforce geometric consistency, and a Geometric-Context Gating (GC-Gate) mechanism. This mechanism functions as a geometrically modulated aperture to dynamically calibrate the fusion of spatial and angular manifolds. Experiments on the Ur-banLF benchmark demonstrate that our method outperforms other state-of-the-art (SOTA) methods, achieving 86.55% mIoU on UrbanLF-Real while maintaining an improved balance between accuracy and linear efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- RTFormer: Efficient Design for Real-Time Semantic Segmentation with TransformerJian Wang, Chenhui Gou, Qiman Wu, Haocheng Feng 等NeurIPS 2022 · 被引用 207 次
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu 等ICML 2024 · 被引用 64 次
- Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesYuchen Duan, Weiyun Wang, Zhe Chen, Xizhou Zhu 等ICLR 2025 · 被引用 10 次
相关 Paper
- Epipolar Consistency-based Network for Structure-Aware LF Semantic SegmentationChen Gao, Youfang Lin, Wenbin Wang, Shuo ZhangACM MM 2025
- Combining Implicit-Explicit View Correlation for Light Field Semantic SegmentationRuixuan Cong, Da Yang, Rongshan Chen, Sizhe Wang 等CVPR 2023
- Learning Non-Local Spatial-Angular Correlation for Light Field Image Super-ResolutionZhengyu Liang, Yingqian Wang, Longguang Wang, Jungang Yang 等ICCV 2023 · 被引用 72 次
- Learning Spatial-angular Fusion for Compressive Light Field Imaging in a Cycle-consistent FrameworkXianqiang Lyu, Zhiyu Zhu, Mantang Guo, Jing Jin 等ACM MM 2021 · 被引用 8 次
- Epipolar Consistent Attention Aggregation Network for Unsupervised Light Field Disparity EstimationChen Gao, Shuo Zhang, Youfang LinICCV 2025 · 被引用 3 次
