GarmentGPT: Compositional Garment Pattern Generation via Discrete Latent Tokenization
Fangsheng Weng, Junhao Chen, Xiang Li, Jie Qin, Hanzhong Guo, ShaochunHao, Xiaoguang Han
摘要
Apparel is a fundamental component of human appearance, making garment digitalization critical for digital human creation. However, sewing pattern creation traditionally relies on the intuition and extensive experience of skilled artisans. This manual bottleneck significantly hinders the scalability of digital garment creation. Existing generative approaches either operate as data replicators without intrinsic understanding of garment construction principles (e.g., diffusion models), or struggle with low-level regression of raw floating-point coordinates (e.g., Vision-Language Models). We present GarmentGPT, the first framework to operationalize latent space generation for sewing patterns. Our approach introduces a novel pipeline where a RVQ-VAE tokenizes continuous pattern boundary curves into discrete codebook indices. A fine-tuned Vision-Language Model then autoregressively predicts these discrete token sequences instead of regressing coordinates, enabling high-level compositional reasoning. This paradigm shift aligns generation with the knowledge-driven, symbolic reasoning capabilities of large language models. To address the data bottleneck for real-world applications, we develop a Data Curation Pipeline that synthesizes over one million photorealistic images paired with GarmentCode, and establish the Real-Garments Benchmark for comprehensive evaluation. Experiments demonstrate that GarmentGPT significantly outperforms existing methods on structured datasets (95.62% Panel Accuracy, 81.84% Stitch Accuracy), validating our discrete compositional paradigm's advantages. Code is available at https://github.com/ChimerAI-MMLab/Garment-GPT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video SynthesisMingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee 等CVPR 2026 · 被引用 13 次
- LottieGPT: Tokenizing Vector Animation for Autoregressive GenerationJunhao Chen, Kejun Gao, Yuehan Cui, Mingze Sun 等CVPR 2026 · 被引用 10 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
相关 Paper
- DressCode: Autoregressively Sewing and Generating Garments from Text GuidanceKai He, Kaixin Yao, Qixuan Zhang, Jingyi Yu 等SIGGRAPH 2024 · 被引用 43 次
- SwiftTailor: Efficient 3D Garment Generation with Geometry Image RepresentationPhuc Pham, Uy Dieu Tran, Binh-Son Hua, Phong NguyenCVPR 2026
- Learning Sewing Patterns via Latent Flow Matching of Implicit FieldsCong Cao, Ren Li, Corentin Dumery, Hao LiSIGGRAPH 2026
- Multimodal Latent Diffusion Model for Complex Sewing Pattern GenerationShengqi Liu, Yuhao Cheng, Zhuo Chen, Xingyu Ren 等ICCV 2025 · 被引用 5 次
- PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D GarmentsZhenyang Li, Lutao Jiang, Yizhou Zhao, Ying-Cong Chen 等SIGGRAPH 2026
