PerLA: Perceptive 3D Language Assistant
Guofeng Mei, Wei Lin, Luigi Riz, Yujiao Wu, Fabio Poiesi, Yiming Wang
Abstract
What is the object on the right side of a gray chair, on top of the table? What is this object? [bbox ] or . radiator black computer monitor SOTA PerLA this is a black suitcase. it is on the floor. the pillow is on the left side of the bed. the pillow is black. SOTA PerLA CLICK CLICK Figure 1. PerLA is a 3D language assistant that integrates local details with global context to learn informative representations of 3D scenes, whereas state-of-the-art (SOTA) 3DLAs focus solely on global context information. PerLA can provide more accurate responses, correctly distinguishing between objects such as a "black computer monitor" and a "black suitcase," where SOTA models instead fail with hallucinated responses. Examples in figures show cases where capturing details from the point cloud matters for accurate output captions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00ba3e93-cc12-4511-86b5-abc3c17ab1bfCited by top-tier papers5
- Ross3d: Reconstructive Visual Instruction Tuning With 3D-AwarenessHaochen Wang, Yucheng Zhao, Tiancai Wang, Haoqiang Fan et al.ICCV 2025 · 7 citations
- Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language ModelsJinlong Li, Liyuan Jiang, Haonan Zhang, Nicu SebeCVPR 2026 · 5 citations
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal InconsistenciesLukas Selch, Yufang Hou, Muhammad Jehanzeb Mirza, Sivan Doveh et al.ICLR 2026 · 2 citations
- Efficient Encoder-Free Fourier-based 3D Large Multimodal ModelGuofeng Mei, Wei Lin, Luigi Riz, Yujiao Wu et al.CVPR 2026 · 2 citations
- MORE-STEM: Long-Short MemOry REcall and Spatio-TEmporal Consistency Model for Query-Driven 3D/4D Point Cloud SegmentationChade Li, Haida Feng, Pengju Zhang, Yihong WuCVPR 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
Related papers
- Video Spatial Reasoning with Object-Centric 3D RolloutHaoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang et al.AAAI 2026 · 3 citations
- SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual GroundingRong Li, Shijie Li, Lingdong Kong, Xulei Yang et al.CVPR 2025
- ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction GenerationLing-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu et al.CVPR 2025
- Towards Learning to Complete Anything in LidarAyça Takmaz, Cristiano Saltori, Neehar Peri, Tim Meinhardt et al.ICML 2025
- CheXplain: Enabling Physicians to Explore and Understand Data-Driven, AI-Enabled Medical Imaging AnalysisYao Xie, Melody Chen, David Kao, Ge Gao et al.CHI 2020 · 137 citations
