SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
Jierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun, Arpit Sahni, Aarush Gupta, Anujraaj Goyal, Dishani Lahiri, Rajesh Singh, Yerlan Idelbayev, Junli Cao, Yanyu Li
2025年份
10顶会引用
摘要
3 HKUST 4 MBZUAI * Equal contribution † Equal advising Project Page: https://snap-research.github.io/snapgen "…, young sudanese female, glamour, natural, front view, extreme detailed and texture skin, …" "a dolphin in an astronaut suit, Animals, Simple Detail" "an old raccoon wearing a top hat and holding an apple, oil painting in the style of van gogh, …" "… llama wearing sunglasses standing on the deck of a spaceship with the Earth in the background, …"
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- PocketSR: The Super-Resolution Expert in Your Pocket MobilesHaoze Sun, Linfeng Jiang, Fan Li, Renjing Pei 等NeurIPS 2025 · 被引用 8 次
- SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationJunsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu 等ICCV 2025 · 被引用 6 次
- HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion ModelsYoung D. Kwon, Rui Li, Sijia Li, Da Li 等AAAI 2026 · 被引用 5 次
- Mobile-VTON: High-Fidelity On-Device Virtual Try-OnZhenchen Wan, Ce Chen, Runqi Lin, Jiaxin Huang 等CVPR 2026 · 被引用 4 次
- NanoSD: Edge Efficient Foundation Model for Real Time Image RestorationSubhajit Sanyal, Srinivas Soumitri Miriyala, Akshay Janardan Bankar, Manjunath Arveti 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Guided Score identity Distillation for Data-Free One-Step Text-to-Image GenerationMingyuan Zhou, Zhendong Wang, Huangjie Zheng, Hai HuangICLR 2025
- Multi-Concept Customization of Text-to-Image DiffusionNupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman 等CVPR 2023
- Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion ModelsAndreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn 等CVPR 2023
- CapHuman: Capture Your Moments in Parallel UniversesChao Liang, Fan Ma, Linchao Zhu, Yingying Deng 等CVPR 2024
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
