RomanTex: Decoupling 3D-Aware Rotary Positional Embedded Multi-Attention Network for Texture Synthesis
Yifei Feng, Mingxin Yang, Shuhui Yang, Sheng Zhang, Jiaao Yu, Zibo Zhao, Yuhong Liu, Jie Jiang, Chunchao Guo
Abstract
Painting textures for existing geometries is a critical yet labor-intensive process in 3D asset generation. Recent advancements in text-to-image (T2I) models have led to significant progress in texture generation. Most existing research approaches this task by first generating images in 2D spaces using image diffusion models, followed by a texture baking process to achieve UV texture. However, these methods often struggle to produce high-quality textures due to inconsistencies among the generated multi-view images, resulting in seams and ghosting artifacts. In contrast, 3D-based texture synthesis methods aim to address these inconsistencies, but they often neglect diffusion model priors, making them challenging to apply to real-world objects. To overcome these limitations, we propose RomanTex, a multiview-based texture generation framework that integrates a multiattention network with an underlying 3D representation, facilitated by our novel 3D-aware Rotary Positional Embedding. Additionally, we incorporate a decoupling characteristic in the multi-attention block to enhance the model's robustness in image-to-texture task, enabling semanticallycorrect back-view synthesis. Furthermore, we introduce a geometry-related Classifier-Free Guidance (CFG) mechanism to further improve the alignment with both geometries and images. Quantitative and qualitative evaluations, along with comprehensive user studies, demonstrate that our method achieves state-of-the-art results in texture quality and consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58d2e70b-d782-4065-8906-fa9b17cdd907Cited by top-tier papers5
- NaTex: Seamless Texture Generation as Latent Color DiffusionZeqiang Lai, Yunfei Zhao, Zibo Zhao, Xin Yang et al.CVPR 2026 · 11 citations
- CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture GenerationChenyu Liu, Hongze CHEN, Jingzhi Bao, Lingting Zhu et al.CVPR 2026 · 3 citations
- MatMart: Material Reconstruction of 3D Objects via DiffusionXiuchao Wu, Pengfei Zhu, Jiangjing Lyu, Xinguo Liu et al.CVPR 2026 · 2 citations
- Toward Richer Material Generation via Procedural Data EnhancementYunchen Yu, Jacob Munkberg, Jon Hasselgren, Chris Cummings et al.SIGGRAPH 2026
- MV2UV: Generating High-quality UV Texture Maps with Multiview PromptsZheng Zhang, Qinchuan Zhang, Yuteng Ye, Zhi Chen et al.CVPR 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- AlignTex: Pixel-Precise Texture Generation from Multi-view ArtworkYuqing Zhang, Hao Xu, Yiqian Wu, Sirui Chen et al.SIGGRAPH 2025 · 4 citations
- MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3DWei Cheng, Juncheng Mu, Xianfang Zeng, Xin Chen et al.CVPR 2025
- UniTEX: Universal High Fidelity Generative Texturing for 3D ShapesYixun Liang, Kunming Luo, Xiao Chen, Rui Chen et al.CVPR 2026 · 27 citations
- Text2Tex: Text-driven Texture Synthesis via Diffusion ModelsDave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov et al.ICCV 2023 · 262 citations
- PixTex: Consistent 3D Texturing via Pixel-Space Multi-View DiffusionYuqing Zhang, Yan-Pei Cao, Hao Xu, Yiqian Wu et al.SIGGRAPH 2026
