AffIn-Space: Learning Affine-Invariant Representations for 3D Spatial Understanding with MLLMs
Zhenyu Lu, Liupeng Li, Jinpeng Wang, Haoqian Kang, Manyuan Zhang, Yan Feng, Ke Chen, Yaowei Wang
Abstract
While MLLMs show promising capacity on general visual understanding, they suffer from geometric fragility : standard visual representations often degrade rapidly under changes in viewpoint and viewing distance. Our analysis identifies that existing paradigms, whether relying on input-level fusion or latent reconstruction, remain entangled with the view-dependent pixel grid, failing to decouple intrinsic 3D structure from extrinsic camera pose. To address this, we introduce AffIn-Space , a framework that enforces strict affine invariance to enable robust spatial understanding. Unlike implicit learning approaches, AffIn-Space introduces a two-stage explicit decoupling mechanism. First, it employs explicit geometric resampling by utilizing decomposed affine quantities (derived from pose features) to spatially align 3D features to a canonical state before fusion. Second, within the MLLM, we implement affine-invariant constraints via an orthogonal projection mechanism, which mathematically strips away pose-dependent noise from the hidden states while retaining recoverable geometric semantics through conditional reconstruction. Extensive experiments on VSI-Bench, ScanQA, SQA3D, Scan2Cap, and EmbodiedScan demonstrate that AffIn-Space achieves state-of-the-art performance. Code and detailed instructions will be publicly released. Crucially, our approach exhibits superior stability against affine perturbations, validating the effectiveness of explicitly modeling geometric invariance for complex spatial tasks. Code will be made available. Extensive experiments show that AffIn-Space achieves state-of-the-art performance on spatial reasoning tasks (VSI-Bench, SQA3D and Scan2Cap), and on spatial grounding tasks (ScanRefer and EmbodiedScan), demonstrating the effectiveness of affine invariant representations for complex spatial understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9344e4e4-b917-4f6a-8ff9-54808147ee2aBuilds on22
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang et al.ICLR 2026 · 318 citations
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 245 citations
- Chat-Scene: Bridging 3D Scene and Large Language Models with Object IdentifiersHaifeng Huang, Yilun Chen, Zehan Wang, Rongjie Huang et al.NeurIPS 2024 · 230 citations
- ScanQA: 3D Question Answering for Spatial Scene UnderstandingDaichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki KawanabeCVPR 2022 · 135 citations
Related papers
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language ModelsRuosen Zhao, Zhikang Zhang, Jialei Xu, Jiahao Chang et al.CVPR 2026 · 21 citations
- G^2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial ReasoningWenbo hu, JINGLI LIN, Yilin Long, Yunlong Ran et al.CVPR 2026
- S^2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural GuidanceBeining Xu, Siting Zhu, Zhao Jin, Junxian Li et al.CVPR 2026
- On the Generalization Capacities of MLLMs for Spatial IntelligenceGongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang et al.ICLR 2026 · 10 citations
- Thinking with Geometry: Active Geometry Integration for Spatial ReasoningHaoyuan Li, Qihang Cao, Tao Tang, Kun Xiang et al.ICML 2026 · 12 citations
