HarmonyView: Harmonizing Consistency and Diversity in One-Image-to-3D
Sangmin Woo, Byeongjun Park, Hyojun Go, Jin-Young Kim, Changick Kim
摘要
Recent progress in single-image 3D generation highlights the importance of multi-view coherency, leveraging 3D priors from large-scale diffusion models pretrained on Internet-scale images. However, the aspect of novel-view diversity remains underexplored within the research landscape due to the ambiguity in converting a 2D image into 3D content, where numerous potential shapes can emerge. Here, we aim to address this research gap by simultaneously addressing both consistency and diversity. Yet, striking a balance between these two aspects poses a considerable challenge due to their inherent trade-offs. This work introduces HarmonyView, a simple yet effective diffusion sampling technique adept at decomposing two intricate aspects in single-image 3D generation: consistency and diversity. This approach paves the way for a more nuanced exploration of the two critical dimensions within the sampling process. Moreover, we propose a new evaluation metric based on CLIP image and text encoders to comprehensively assess the diversity of the generated views, which closely aligns with human evaluators' judgments. In experiments, HarmonyView achieves a harmonious balance, demonstrating a win-win scenario in both consistency and diversity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian GenerationChubin Zhang, Hongliang Song, Yi Wei, Chen Yu 等NeurIPS 2024 · 被引用 40 次
- Denoising Task Routing for Diffusion ModelsByeongjun Park, Sangmin Woo, Hyojun Go, Jin-Young Kim 等ICLR 2024 · 被引用 26 次
- MVD^2: Efficient Multiview 3D Reconstruction for Multiview DiffusionXin-Yang Zheng, Hao Pan, Yu-Xiao Guo, Xin Tong 等SIGGRAPH 2024 · 被引用 12 次
- ReDirector: Creating Any-Length Video Retakes with Rotary Camera EncodingByeongjun Park, Byung-Hoon Kim, Hyungjin Chung, Jong ChulCVPR 2026 · 被引用 10 次
- Diffusion Model Patching via Mixture-of-PromptsSeokil Ham, Sangmin Woo, Jin-Young Kim, Hyojun Go 等AAAI 2025 · 被引用 9 次
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware DiffusionHaoran Wei, Wencheng Han, Xingping Dong, Jianbing ShenAAAI 2026
- SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageYuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long 等ICLR 2024 · 被引用 685 次
- Vistadream: Sampling Multiview Consistent Images for Single-View Scene ReconstructionHaiping Wang, Yuan Liu, Ziwei Liu, Wenping Wang 等ICCV 2025 · 被引用 8 次
- Novel View Synthesis with Diffusion ModelsDaniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho 等ICLR 2023 · 被引用 63 次
- WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single ImageJiwoo Park, Tae Eun Choi, Youngjun Jun, Seong Jae HwangICCV 2025
