FISHuman: Fine-grained Single-image 3D Human Reconstruction via Multi-view 4D Remeshing
Hanxi Liu, Yifang Men, Zhouhui Lian
Abstract
Single-image 3D human reconstruction holds significant promise due to its convenience and high demand in various applications. Previous methods have garnered tremendous progress by employing 2D multi-view diffusion models to generate auxiliary views as reconstruction priors, but they struggle with 3D inconsistencies and limited generalization capabilities. In this paper, we present FISHuman, which aims to generate fine-grained, high-fidelity, and content-wise diverse 3D humans from a single-view input, providing production-ready 3D assets. We propose an elaborately designed workflow that reconstructs dynamic 3D meshes from multi-view inconsistent guidance. Specifically, we adapt a dual-stream transformer-based video diffusion model to generate cross-modally aligned multi-view RGB and normal sequences. We find that naively employing static 3D reconstruction can lead to geometric distortions and texture blurriness, due to the lack of 3D awareness within the generated frames. To address this, we introduce a novel 4D remeshing module that explicitly disentangles the learning of the globally shared canonical mesh and transient variations by tracking per-vertex deformations under different viewpoints. The topological consistency of the deformed meshes inherently enables the optimization of a unified UV representation that effectively integrates appearance attributes across frames. Both qualitative and quantitative experimental results demonstrate the superiority of our method over prior works in terms of appearance realism, geometric fineness, and generalization diversity. We also showcase the applicability of our reconstructed avatars for downstream applications including animation and 3D editing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 456d9aeb-b7ea-47ce-9bb9-ef35e7f62fecBuilds on55
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai et al.ICLR 2024 · 973 citations
Related papers
- GAS: Generative Avatar Synthesis from a Single ImageYixing Lu, Junting Dong, Youngjoong Kwon, Qin Zhao et al.ICCV 2025 · 5 citations
- MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative RefinementXu He, Zhiyong Wu, Xiaoyu Li, Di Kang et al.AAAI 2025 · 11 citations
- GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human DataWentao Wang, Hang Ye, Fangzhou Hong, Xue Yang et al.NeurIPS 2025 · 6 citations
- Intrinsic Geometry-Appearance Consistency Optimization for Sparse-View Gaussian SplattingKaiqiang Xiong, Rui Peng, Jiahao Wu, Zhanke Wang et al.CVPR 2026 · 2 citations
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human ReconstructionWenyue Chen, Peng Li, Wangguandong Zheng, Chengfeng Zhao et al.NeurIPS 2025 · 8 citations
