Struct-Align: Zero-Shot Text-to-3D Scene Retrieval via Locality-Aware Structural Alignment
Xiong Li, Yikang Yan, Zhenyu Wen, Jie Su, Qi Chen, Zhen Hong
Abstract
Text-to-3D Scene Retrieval (T3SR) aims to retrieve 3D scenes that match users' linguistic queries, enabling intuitive access to 3D scene repositories. Existing approaches rely on joint embedding learning with large amounts of paired text–scene data, which is expensive to collect and often fails to generalize under open-vocabulary queries and diverse scene distributions. In this paper, we propose Struct-Align, a foundation-model-driven framework for zero-shot T3SR that eliminates the need for paired training data. Our key insight is to reformulate T3SR as a single-modality structural alignment problem by converting both 3D scenes and textual queries into a shared, schema-aligned textual representation compatible with pretrained text embedding models. To reliably derive such representations from complex 3D environments, we introduce a role-decomposed scene structuring pipeline that mitigates generative instability and produces semantically consistent scene depictions. To address the inherent semantic asymmetry between query and scene representations, we further propose a locality-aware structural matching strategy that explicitly localizes query intent and performs instance- and relation-level alignment within query-relevant substructures. Extensive experiments on multiple benchmarks demonstrate that Struct-Align outperforms both training-based and zero-shot baselines while exhibiting strong robustness to domain shift.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f5fee10a-2098-4a1a-b8c5-8cb30507dc02Related papers
- Interactive Cross-modal Learning for Text-3D Scene RetrievalYanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin et al.NeurIPS 2025 · 9 citations
- Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and GroundingYanglin Feng, Hongyuan Zhu, Dezhong Peng, Xi Peng et al.NeurIPS 2025 · 6 citations
- OpenScene: 3D Scene Understanding with Open VocabulariesSongyou Peng, Kyle Genova, Chiyu Max Jiang, Andrea Tagliasacchi et al.CVPR 2023
- Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set RelationshipsSebastian Koch, Narunas Vaskevicius, Mirco Colosi, Pedro Hermosilla et al.CVPR 2024
- OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned ImagesYe Mao, Junpeng Jing, Krystian MikolajczykNeurIPS 2024 · 10 citations
