UMFuse: Unified Multi View Fusion for Human Editing applications
Rishabh Jain, Mayur Hemani, Duygu Ceylan, Krishna Kumar Singh, Jingwan Lu, Mausoom Sarkar, Balaji Krishnamurthy
Abstract
Numerous pose-guided human editing methods have been explored by the vision community due to their extensive practical applications. However, most of these methods still use an image-to-image formulation in which a single image is given as input to produce an edited image as output. This objective becomes ill-defined in cases when the target pose differs significantly from the input pose. Existing methods then resort to in-painting or style transfer to handle occlusions and preserve content. In this paper, we explore the utilization of multiple views to minimize the issue of missing information and generate an accurate representation of the underlying human model. To fuse knowledge from multiple viewpoints, we design a multi-view fusion network that takes the pose key points and texture from multiple source images and generates an explainable per-pixel appearance retrieval map. Thereafter, the encodings from a separate network (trained on a single-view human reposing task) are merged in the latent space. This enables us to generate accurate, precise, and visually coherent images for different editing tasks. We show the application of our network on two newly proposed tasks - Multi-view human reposing and Mix&Match Human Image generation. Additionally, we study the limitations of single-view editing and scenarios in which multi-view provides a better alternative. Datasplits and results can be found at Project Webpage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8982809f-8b70-4ae2-8de9-bc79a7ed414eBuilds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular VideoChung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron et al.CVPR 2022 · 411 citations
- Dressing in Order: Recurrent Person Image Generation for Pose Transfer, Virtual Try-on and Outfit EditingAiyu Cui, Daniel McKee, Svetlana LazebnikICCV 2021 · 110 citations
- Neural Texture Extraction and Distribution for Controllable Person Image SynthesisYurui Ren, Xiaoqing Fan, Ge Li, Shan Liu et al.CVPR 2022 · 81 citations
Related papers
- IMAGPose: A Unified Conditional Framework for Pose-Guided Person GenerationFei Shen, Jinhui TangNeurIPS 2024 · 172 citations
- UniHuman: A Unified Model For Editing Human Images in the WildNannan Li, Qing Liu, Krishna Kumar Singh, Yilin Wang et al.CVPR 2024
- VGFlow: Visibility guided Flow Network for Human ReposingRishabh Jain, Krishna Kumar Singh, Mayur Hemani, Jingwan Lu et al.CVPR 2023
- MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image GenerationTianxiang Ma, Bo Peng, Wei Wang, Jing DongCVPR 2021
- Pose Guided Image Generation from Misaligned Sources via Residual Flow Based CorrectionJiawei Lu, He Wang, Tianjia Shao, Yin Yang et al.AAAI 2022 · 5 citations
