Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
Hyojun Go, Dominik Narnhofer, Goutam Bhat, Prune Truong, Federico Tombari, Konrad Schindler
Abstract
The rapid progress of large, pretrained models for both visual content generation and 3D reconstruction opens up new possibilities for text-to-3D generation. Intuitively, one could obtain a formidable 3D scene generator if one were able to combine the power of a modern latent text-to-video model as "generator" with the geometric abilities of a recent (feedforward) 3D reconstruction system as "decoder". We introduce VIST3A, a general framework that does just that, addressing two main challenges. First, the two components must be joined in a way that preserves the rich knowledge encoded in their weights. We revisit model stitching, i.e., we identify the layer in the 3D decoder that best matches the latent representation produced by the text-to-video generator and stitch the two parts together. That operation requires only a small dataset and no labels. Second, the text-to-video generator must be aligned with the stitched 3D decoder, to ensure that the generated latents are decodable into consistent, perceptually convincing 3D scene geometry. To that end, we adapt direct reward finetuning, a popular technique for human preference alignment. We evaluate the proposed VIST3A approach with different video generators and 3D reconstruction models. All tested pairings markedly improve over prior text-to-3D models that output Gaussian splats. Moreover, by choosing a suitable 3D base model, VIST3A also enables high-quality text-to-pointmap generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96f5a37b-d788-43c9-849b-f08fb49e5905Cited by top-tier papers2
- Understanding, Accelerating, and Improving MeanFlow TrainingJin-Young Kim, Hyojun Go, Lea Bogensperger, Julius Erbach et al.CVPR 2026 · 4 citations
- Splatent: Splatting Diffusion Latents for Novel View SynthesisOr Hirschorn, Omer Sela, Inbar Huberman-Spiegelglas, Netalee Efrat Sela et al.CVPR 2026 · 2 citations
Builds on74
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose EstimationSiyou Lin, Zhou Xue, Hongwen Zhang, Liang An et al.SIGGRAPH 2026
- Gen3R: 3D Scene Generation Meets Feed-Forward ReconstructionJiaxin Huang, Yuanbo Yang, Bangbang Yang, Lin Ma et al.CVPR 2026 · 24 citations
- VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint ModelingHyojun Go, Byeongjun Park, Hyelin Nam, Byung-Hoon Kim et al.ICCV 2025 · 1 citation
- DreamCS: Geometry-Aware Text-to-3D Generation with Unpaired 3D Reward SupervisionXiandong Zou, Ruihao Xia, Hongsong Wang, Pan ZhouICLR 2026 · 6 citations
- SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric SteeringByeongjun Park, Hyojun Go, Hyelin Nam, Byung-Hoon Kim et al.ICCV 2025 · 1 citation
