GLIGEN: Open-Set Grounded Text-to-Image Generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, Yong Jae Lee
Abstract
https://gligen.github.io/ Caption: "A woman sitting in a restaurant with a pizza in front of her " Grounded text: table, pizza, person, wall, car, paper, chair, window, bottle, cup Caption: "a baby girl / monkey / Hormer Simpson / is scratching her/its head" Grounded keypoints: plotted dots on the left image Caption: "A dog / bird / helmet / backpack is on the grass" Grounded image: red inset Caption: "Elon Musk and Emma Watson on a movie poster" Grounded text: Elon Musk, Emma Watson; Grounded style image: blue inset Caption: "A vibrant colorful bird sitting on tree branch" Grounded depth map: the left image Caption: "A young boy with white powder on his face looks away" Grounded HED map: the left image Caption: "Cars park on the snowy street" Grounded normal map: the left image Caption: "A living room filled with lots of furniture and plants" Grounded semantic map: the left image § Part of the work performed at Microsoft; ¶ Co-senior authors This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3ed7795-5f94-4509-a2b0-a7b0c240f37cCited by top-tier papers397
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsShihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao et al.NeurIPS 2023 · 505 citations
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
- BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained DiffusionJinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu et al.ICCV 2023 · 313 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- PrEditor3D: Fast and Precise 3D Shape EditingZiya Erkoç, Can Gümeli, Chaoyang Wang, Matthias Nießner et al.CVPR 2025
- Multi-Concept Customization of Text-to-Image DiffusionNupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman et al.CVPR 2023
- 3DEnhancer: Consistent Multi-View Diffusion for 3D EnhancementYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan et al.CVPR 2025
- HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled GenerationBoyuan Wang, Xiaofeng Wang, Chaojun Ni, Guosheng Zhao et al.CVPR 2025
- Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern AnalysisMang Tik Chiu, Xingqian Xu, Yunchao Wei, Zilong Huang et al.CVPR 2020
