S2-Edit3DV: Diffusion-Guided Style Meets Structure for Consistent Multi-View 3D Video Generation
Yuqi Chen, Xiubo Liang, Yu Zhao, Hongzhi Wang, Weidong Geng
Abstract
Consistently stylizing and editing 3D objects from multiple viewpoints is crucial for immersive applications such as virtual reality, augmented reality, and digital entertainment. Nevertheless, existing methods frequently face significant challenges, including inconsistent textures, pronounced drifting artifacts, and compromised geometric integrity when rendered from various perspectives. To effectively address these limitations, we introduce S2-Edit3DV, a novel diffusion-guided framework that reframes multi-view 3D objects editing as a temporally coherent video editing problem. By exploiting the robust single-view generative capabilities of SV3D, our approach reliably propagates initial style edits across different viewpoints, substantially mitigating drifting artifacts prevalent in current video-based editing methods. To further enhance semantic precision and structural preservation, we propose two innovative techniques: Attention-based Differential Style Injection (ADSI) and Adaptive Structural-aware Plug-and-Play (AS-PnP). ADSI utilizes attention-driven semantic embeddings for adaptive and precise style injection, effectively reducing semantic hallucinations. AS-PnP strategically modulates stylized latent features, balancing artistic expression with strict structural coherence. Comprehensive evaluations and ablation studies demonstrate that our proposed framework significantly enhances multi-view consistency, preserves fine-grained geometric details, and ensures accurate semantic alignment, showcasing superior performance and practical value for generating high-quality, creatively stylized, and structurally robust objects.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Edit360: 2D Image Edits to 3D Assets From Any AngleJunchao Huang, Xinting Hu, Shaoshuai Shi, Zhuotao Tian et al.ICCV 2025 · 2 citations
- Diffusion Feature Field for Text-based 3D Editing with Gaussian SplattingEunseo Koh, Sangeek Hyun, MinKyu Lee, Jiwoo Chung et al.NeurIPS 2025 · 5 citations
- DiffStyle3D: Consistent 3D Gaussian Stylization via Attention OptimizationYitong Yang, Yinglin Wang, Xuexin Liu, Jing Wang et al.ICML 2026 · 2 citations
- Scene-Level Appearance Transfer with Semantic CorrespondencesLiyuan Zhu, Shengqu Cai, Shengyu Huang, Gordon Wetzstein et al.SIGGRAPH 2025 · 3 citations
- SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D EditingRuihuang Li, Liyi Chen, Zhengqiang Zhang, Varun Jampani et al.AAAI 2025 · 4 citations
