iVS-Net: Learning Human View Synthesis from Internet Videos
Junting Dong, Qi Fang, Tianshuo Yang, Qing Shuai, Chengyu Qiao, Sida Peng
Abstract
Recent advances in implicit neural representations make it possible to generate free-viewpoint videos of the human from sparse view images. To avoid the expensive training for each person, previous methods adopt the generalizable human model and demonstrate impressive results. However, these methods usually rely on limited multi-view images typically collected in the studio or commercial high-quality 3D scans for training, which heavily prohibits their generalization capability for in-the-wild images. To solve this problem, we propose a new approach to learn a generalizable human model from a new source of data, i.e., Internet videos. These videos capture various human appearances and poses and record the performers from abundant viewpoints. To exploit the Internet data, we present a video self-supervised pipeline to enforce the local appearance consistency of each body part over different frames of the same video. Once learned, the human model enables realistic novel view synthesis from a single input image. Experiments show that our method can generate high-quality view synthesis on in-the-wild images while only training on monocular videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0e92c5a-ceab-4d9e-9774-ec96b19f37faBuilds on24
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
- Volume Rendering of Neural Implicit SurfacesLior Yariv, Jiatao Gu, Yoni Kasten, Yaron LipmanNeurIPS 2021 · 1,421 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- PlenOctrees for Real-time Rendering of Neural Radiance FieldsAlex Yu, Ruilong Li, Matthew Tancik, Hao Li et al.ICCV 2021 · 1,284 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
Related papers
- Neural Body: Implicit Neural Representations With Structured Latent Codes for Novel View Synthesis of Dynamic HumansSida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang et al.CVPR 2021
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and PoseShih-Yang Su, Frank Yu, Michael Zollhöfer, Helge RhodinNeurIPS 2021 · 316 citations
- Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance VideosYasamin Jafarian, Hyun Soo ParkCVPR 2021
- GM-NeRF: Learning Generalizable Model-Based Neural Radiance Fields from Multi-View ImagesJianchuan Chen, Wentao Yi, Liqian Ma, Xu Jia et al.CVPR 2023
- TotalSelfScan: Learning Full-body Avatars from Self-Portrait Videos of Faces, Hands, and BodiesJunting Dong, Qi Fang, Yudong Guo, Sida Peng et al.NeurIPS 2022 · 24 citations
