MatLat: Material Latent Space for PBR Texture Generation
Kyeongmin Yeo, Yunhong Min, Jaihoon Kim, Minhyuk Sung
Abstract
We propose a generative framework for producing high-quality PBR textures on a given 3D mesh. As large-scale PBR texture datasets are scarce, our approach focuses on effectively leveraging the embedding space and diffusion priors of pretrained latent image generative models while learning a material latent space, MatLat , through targeted fine-tuning. Unlike prior methods that freeze the embedding network and thus lead to distribution shifts when encoding additional PBR channels and hinder subsequent diffusion training, we fine-tune the pretrained VAE so that new material channels can be incorporated with minimal latent distribution deviation. We further show that correspondence-aware attention alone is insufficient for cross-view consistency unless the latent-to-image mapping preserves locality. To enforce this locality, we introduce a regularization in the VAE fine-tuning that crops latent patches, decodes them, and aligns the corresponding image regions to maintain strong pixel–latent spatial correspondence. Ablations studies and comparison with previous baselines demonstrate that our framework improves PBR texture fidelity and that each component is critical for achieving state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 505f021b-20ea-4369-b5af-610bd14a649eBuilds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- MaterialMVP: Illumination-Invariant Material Generation via Multi-View PBR DiffusionZebin He, Mingxin Yang, Shuhui Yang, Yixuan Tang et al.ICCV 2025 · 3 citations
- PixTex: Consistent 3D Texturing via Pixel-Space Multi-View DiffusionYuqing Zhang, Yan-Pei Cao, Hao Xu, Yiqian Wu et al.SIGGRAPH 2026
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationZilong Chen, Yikai Wang, Wenqiang Sun, Feng Wang et al.CVPR 2025
- NaTex: Seamless Texture Generation as Latent Color DiffusionZeqiang Lai, Yunfei Zhao, Zibo Zhao, Xin Yang et al.CVPR 2026 · 11 citations
- GenesisTex2: Stable, Consistent and High-Quality Text-to-Texture GenerationJiawei Lu, Yingpeng Zhang, Zengjun Zhao, He Wang et al.AAAI 2025 · 10 citations
