Voxify3D: Pixel Art Meets Volumetric Rendering
Yi-Chuan Huang, Jiewen Chan, Hao-Jen Chien, Yu-Lun Liu
Abstract
Voxel art is a distinctive stylization widely used in games and digital media, yet automated generation from 3D meshes remains challenging due to conflicting requirements of geometric abstraction, semantic preservation, and discrete color coherence. Existing methods either over-simplify geometry or fail to achieve the pixel-precise, palette-constrained aesthetics of voxel art. We introduce Voxify3D, a differentiable two-stage framework bridging 3D mesh optimization with 2D pixel art supervision. Our core innovation lies in the synergistic integration of three components: (1) orthographic pixel art supervision that eliminates perspective distortion for precise voxel-pixel alignment; (2) patch-based CLIP alignment that preserves semantics across discretization levels; (3) palette-constrained Gumbel-Softmax quantization enabling differentiable optimization over discrete color spaces with controllable palette strategies. This integration addresses fundamental challenges: semantic preservation under extreme discretization, pixel-art aesthetics through volumetric rendering, and end-to-end discrete optimization. Experiments show superior performance (37.12 CLIP-IQA, 77.90% user preference) across diverse characters and controllable abstraction (2-8 colors, 20x-50x resolutions). Project page: https://yichuanh.github.io/Voxify-3D/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60028c24-eff4-45c7-8539-fd2d4b191c73Builds on63
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Plenoxels: Radiance Fields without Neural NetworksSara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen et al.CVPR 2022 · 1,237 citations
Related papers
- Unsupervised Representation Learning for 3D Mesh Parameterization with Semantic and Visibility ObjectivesAmirHossein Zamani, Bruno Roy, Arianna RampiniICLR 2026 · 1 citation
- 3DStyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with 2D Diffusion ModelsHaibo Yang, Yang Chen, Yingwei Pan, Ting Yao et al.ACM MM 2023 · 23 citations
- AlignTex: Pixel-Precise Texture Generation from Multi-view ArtworkYuqing Zhang, Hao Xu, Yiqian Wu, Sirui Chen et al.SIGGRAPH 2025 · 4 citations
- VoxGRAF: Fast 3D-Aware Image Synthesis with Sparse Voxel GridsKatja Schwarz, Axel Sauer, Michael Niemeyer, Yiyi Liao et al.NeurIPS 2022 · 185 citations
- Convolutional Generation of Textured 3D MeshesDario Pavllo, Graham Spinks, Thomas Hofmann, Marie-Francine Moens et al.NeurIPS 2020 · 71 citations
