TVDRNet: Text-driven Viewpoint Optimization via Differentiable Rendering for 3D Reasoning Segmentation
Tingran Wang, Changshuo Wang, Pinjie Xu, ZhangHuang, Yuan Shi, Linjun Sun, Weijun Li
摘要
Three-dimensional (3D) reasoning segmentation aims at segmenting target objects based on text instructions and 3D spatial cues. Recent efforts in 3D reasoning leverage Multimodal Large Language Models (MLLMs) to bridge the gap between text and 3D data. However, since MLLMs are primarily trained on text-image pairs, directly adapting them to unstructured 3D point clouds often fails to capture implicit semantic intent and reliably localize objects. This paper introduces TV-DRNet to address these challenges. Inspired by Active Vision theory, where humans selectively choose optimal viewpoints to better observe targets, TVDRNet employs a differentiable renderer to simulate this active process in 3D perception. By using text instructions as supervision to optimize intrinsic and extrinsic rendering parameters, the TVDRNet identifies the optimal viewpoints for observing the 3D scene, and therefore learning 'where to look' based on what the text instruction 'asked to find'. This process generates informative, task-relevant 2D images that are compatible with MLLMs. TVDRNet comprises: (1) the AVPL module, establishing a learnable mapping from semantics to optimal rendering parameters; and (2) the MGL module, fusing multi-modalities via semantic grouping to guide mask generation. Experiments show TVDRNet achieves the state-of-the-art performance in 3D reasoning segmentation (Reason3D, Instruct3D) and 3D visual grounding (ScanRefer) benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun 等NeurIPS 2020 · 被引用 1,010 次
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 被引用 789 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys 等NeurIPS 2023 · 被引用 389 次
相关 Paper
- MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning SegmentationJiaxin Huang, Runnan Chen, Ziwen Li, Zhengqing Gao 等NeurIPS 2025 · 被引用 18 次
- S^2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural GuidanceBeining Xu, Siting Zhu, Zhao Jin, Junxian Li 等CVPR 2026
- Enhancing Spatial Reasoning in Multimodal Large Language Models Through Reasoning-Based SegmentationZhenhua Ning, Zhuotao Tian, Shaoshuai Shi, Guangming Lu 等ICCV 2025
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing 等AAAI 2026 · 被引用 4 次
- ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and ReasoningZhenyang Liu, Yikai Wang, Sixiao Zheng, Tongying Pan 等CVPR 2025
