PerLA: Perceptive 3D Language Assistant
Guofeng Mei, Wei Lin, Luigi Riz, Yujiao Wu, Fabio Poiesi, Yiming Wang
摘要
What is the object on the right side of a gray chair, on top of the table? What is this object? [bbox ] or . radiator black computer monitor SOTA PerLA this is a black suitcase. it is on the floor. the pillow is on the left side of the bed. the pillow is black. SOTA PerLA CLICK CLICK Figure 1. PerLA is a 3D language assistant that integrates local details with global context to learn informative representations of 3D scenes, whereas state-of-the-art (SOTA) 3DLAs focus solely on global context information. PerLA can provide more accurate responses, correctly distinguishing between objects such as a "black computer monitor" and a "black suitcase," where SOTA models instead fail with hallucinated responses. Examples in figures show cases where capturing details from the point cloud matters for accurate output captions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Ross3d: Reconstructive Visual Instruction Tuning With 3D-AwarenessHaochen Wang, Yucheng Zhao, Tiancai Wang, Haoqiang Fan 等ICCV 2025 · 被引用 7 次
- Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language ModelsJinlong Li, Liyuan Jiang, Haonan Zhang, Nicu SebeCVPR 2026 · 被引用 5 次
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal InconsistenciesLukas Selch, Yufang Hou, Muhammad Jehanzeb Mirza, Sivan Doveh 等ICLR 2026 · 被引用 2 次
- Efficient Encoder-Free Fourier-based 3D Large Multimodal ModelGuofeng Mei, Wei Lin, Luigi Riz, Yujiao Wu 等CVPR 2026 · 被引用 2 次
- MORE-STEM: Long-Short MemOry REcall and Spatio-TEmporal Consistency Model for Query-Driven 3D/4D Point Cloud SegmentationChade Li, Haida Feng, Pengju Zhang, Yihong WuCVPR 2026
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
相关 Paper
- Video Spatial Reasoning with Object-Centric 3D RolloutHaoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang 等AAAI 2026 · 被引用 3 次
- SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual GroundingRong Li, Shijie Li, Lingdong Kong, Xulei Yang 等CVPR 2025
- ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction GenerationLing-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu 等CVPR 2025
- Towards Learning to Complete Anything in LidarAyça Takmaz, Cristiano Saltori, Neehar Peri, Tim Meinhardt 等ICML 2025
- CheXplain: Enabling Physicians to Explore and Understand Data-Driven, AI-Enabled Medical Imaging AnalysisYao Xie, Melody Chen, David Kao, Ge Gao 等CHI 2020 · 被引用 137 次
