MNSRNet: Multimodal Transformer Network for 3D Surface Super-Resolution
Wuyuan Xie, Tengcong Huang, Miaohui Wang
Abstract
With the rapid development of display technology, it has become an urgent need to obtain realistic 3D surfaces with as high-quality as possible. Due to the unstructured and irregular nature of 3D object data, it is usually difficult to obtain high-quality surface details and geometry textures at a low cost. In this article, we propose an effective multimodal-driven deep neural network to perform 3D surface super-resolution in 2D normal domain, which is simple, accurate, and robust to the above difficulty. To leverage the multimodal information from different perspectives, we jointly consider the texture, depth, and normal modalities to simultaneously restore fine-grained surface details as well as preserve geometry structures. To better utilize the cross-modality information, we explore a two-bridge normal method with a transformer structure for feature alignment, and investigate an affine transform module for fusing multimodal features. Extensive experimental results on public and our newly constructed photometric stereo dataset demonstrate that the proposed method delivers promising surface geometry details compared with nine competitive schemes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Just Noticeable Visual Redundancy Forecasting: A Deep Multimodal-Driven ApproachWuyuan Xie, Shukang Wang, Sukun Tian, Lirong Huang et al.AAAI 2023 · 5 citations
- HybridMQA: Exploring Geometry-Texture Interactions for Colored Mesh Quality AssessmentArmin Shafiee Sarvestani, Sheyang Tang, Zhou WangCVPR 2025
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Neural subdivisionHsueh-Ti Derek Liu, Vladimir G. Kim, Siddhartha Chaudhuri, Noam Aigerman et al.SIGGRAPH 2020 · 57 citations
- PU-EVA: An Edge-Vector based Approximation Solution for Flexible-scale Point Cloud UpsamplingLuqing Luo, Lulu Tang, Wanyi Zhou, Shizheng Wang et al.ICCV 2021 · 42 citations
- DualConvMesh-Net: Joint Geodesic and Euclidean Convolutions on 3D MeshesJonas Schult, Francis Engelmann, Theodora Kontogianni, Bastian LeibeCVPR 2020
- Closed-Loop Matters: Dual Regression Networks for Single Image Super-ResolutionYong Guo, Jian Chen, Jingdong Wang, Qi Chen et al.CVPR 2020
Related papers
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent AlignmentTongfan Guan, Jiaxin Guo, Chen Wang, Yun-Hui LiuICCV 2025 · 6 citations
- Uncalibrated Neural Inverse Rendering for Photometric Stereo of General SurfacesBerk Kaya, Suryansh Kumar, Carlos E. P. de Oliveira, Vittorio Ferrari et al.CVPR 2021
- Any Resolution Any Geometry: From Multi-View To Multi-PatchWenqing Cui, Zhenyu Li, Mykola Lavreniuk, Jian Shi et al.CVPR 2026 · 2 citations
- PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface ReconstructionMingzhi Pei, Xu Cao, Xiangyi Wang, Heng Guo et al.CVPR 2025
- SuperNormal: Neural Surface Reconstruction via Multi-View Normal IntegrationXu Cao, Takafumi TaketomiCVPR 2024
