Discrete Latent Perspective Learning for Segmentation and Detection
Deyi Ji, Feng Zhao, Lanyun Zhu, Wenwei Jin, Hongtao Lu, Jieping Ye
Abstract
In this paper, we address the challenge of Perspective-Invariant Learning in machine learning and computer vision, which involves enabling a network to understand images from varying perspectives to achieve consistent semantic interpretation. While standard approaches rely on the labor-intensive collection of multi-view images or limited data augmentation techniques, we propose a novel framework, Discrete Latent Perspective Learning (DLPL), for latent multi-perspective fusion learning using conventional single-view images. DLPL comprises three main modules: Perspective Discrete Decomposition (PDD), Perspective Homography Transformation (PHT), and Perspective Invariant Attention (PIA), which work together to discretize visual features, transform perspectives, and fuse multi-perspective semantic information, respectively. DLPL is a universal perspective learning framework applicable to a variety of scenarios and vision tasks. Extensive experiments demonstrate that DLPL significantly enhances the network's capacity to depict images across diverse scenarios (daily photos, UAV, auto-driving) and tasks (detection, segmentation).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Hybrid Mamba for Few-Shot SegmentationQianxiong Xu, Xuanyi Liu, Lanyun Zhu, Guosheng Lin et al.NeurIPS 2024 · 49 citations
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal RetrievalLanyun Zhu, Deyi Ji, Tianrun Chen, Haiyang Wu et al.NeurIPS 2025 · 12 citations
- Generating Negative Samples for Multi-Modal RecommendationYanbiao Ji, Dan Luo, Chang Liu, Shaokai Wu et al.ACM MM 2025 · 2 citations
- FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component AnalysisJiangtong Tan, Hu Yu, Jie Huang, Jie Xiao et al.CVPR 2025
- SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language ModelingQi Zhu, Jiangwei Lao, Deyi Ji, Junwei Luo et al.CVPR 2025
Builds on16
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
Related papers
- UniFusion: A Unified Image Fusion Framework with Robust Representation and Source-Aware PreservationXingyuan Li, Songcheng Du, Yang Zou, Haoyuan Xu et al.CVPR 2026 · 6 citations
- UniMLVG: Unified Framework for Multi-View Long Video Generation with Comprehensive Control Capabilities for Autonomous DrivingRui Chen, Zehuan Wu, Yichen Liu, Yuxin Guo et al.ICCV 2025 · 2 citations
- Generalizable Multi-Camera 3D Object Detection from a Single Source via Fourier Cross-View LearningXue Zhao, Qinying Gu, Xinbing Wang, Chenghu Zhou et al.ICML 2025
- VISTA: Boosting 3D Object Detection via Dual Cross-VIew SpaTial AttentionShengheng Deng, Zhihao Liang, Lin Sun, Kui JiaCVPR 2022 · 92 citations
- UniDrive: Towards Universal Driving Perception Across Camera ConfigurationsYe Li, Wenzhao Zheng, Xiaonan Huang, Kurt KeutzerICLR 2025
