GLIGEN: Open-Set Grounded Text-to-Image Generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, Yong Jae Lee
摘要
https://gligen.github.io/ Caption: "A woman sitting in a restaurant with a pizza in front of her " Grounded text: table, pizza, person, wall, car, paper, chair, window, bottle, cup Caption: "a baby girl / monkey / Hormer Simpson / is scratching her/its head" Grounded keypoints: plotted dots on the left image Caption: "A dog / bird / helmet / backpack is on the grass" Grounded image: red inset Caption: "Elon Musk and Emma Watson on a movie poster" Grounded text: Elon Musk, Emma Watson; Grounded style image: blue inset Caption: "A vibrant colorful bird sitting on tree branch" Grounded depth map: the left image Caption: "A young boy with white powder on his face looks away" Grounded HED map: the left image Caption: "Cars park on the snowy street" Grounded normal map: the left image Caption: "A living room filled with lots of furniture and plants" Grounded semantic map: the left image § Part of the work performed at Microsoft; ¶ Co-senior authors This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper397
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsShihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao 等NeurIPS 2023 · 被引用 505 次
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani 等NeurIPS 2023 · 被引用 462 次
- BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained DiffusionJinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu 等ICCV 2023 · 被引用 313 次
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- PrEditor3D: Fast and Precise 3D Shape EditingZiya Erkoç, Can Gümeli, Chaoyang Wang, Matthias Nießner 等CVPR 2025
- Multi-Concept Customization of Text-to-Image DiffusionNupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman 等CVPR 2023
- 3DEnhancer: Consistent Multi-View Diffusion for 3D EnhancementYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan 等CVPR 2025
- HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled GenerationBoyuan Wang, Xiaofeng Wang, Chaojun Ni, Guosheng Zhao 等CVPR 2025
- Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern AnalysisMang Tik Chiu, Xingqian Xu, Yunchao Wei, Zilong Huang 等CVPR 2020
