Reject Decoding via Language-Vision Models for Text-to-Image Synthesis
Fuxiang Wu, Liu Liu, Fusheng Hao, Fengxiang He, Lei Wang, Jun Cheng
摘要
Transformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, the common practice is drawing multi-paths from the transformer-based models and re-ranking the multi-images decoded from multi-paths to find the best one and filter out others. Therefore, the computing procedure of excluding images may be inefficient. To improve the effectiveness and efficiency of decoding, we exploit a reject decoding algorithm with tiny multi-modal models to enlarge the searching space and exclude the useless paths as early as possible. Specifically, we build tiny multi-modal models to evaluate the similarities between the partial paths and the caption at multi scales. Then, we propose a reject decoding algorithm to exclude some lowest quality partial paths at the inner steps. Thus, under the same computing load as the original decoding, we could search across more multi-paths to improve the decoding efficiency and synthesizing quality. The experiments conducted on the MS-COCO dataset and large-scale datasets show that the proposed reject decoding algorithm can exclude the useless paths and enlarge the searching paths to improve the synthesizing quality by consuming less time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen 等CVPR 2022 · 被引用 607 次
相关 Paper
- DeeCap: Dynamic Early Exiting for Efficient Image CaptioningZhengcong Fei, Xu Yan, Shuhui Wang, Qi TianCVPR 2022 · 被引用 39 次
- VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-wordsXiaopeng Lu, Tiancheng Zhao, Kyusong LeeACL 2021
- Text-to-Image Synthesis based on Object-Guided Joint-Decoding TransformerFuxiang Wu, Liu Liu, Fusheng Hao, Fengxiang He 等CVPR 2022 · 被引用 13 次
- Semantic-Conditional Diffusion Networks for Image CaptioningJianjie Luo, Yehao Li, Yingwei Pan, Ting Yao 等CVPR 2023
- FlashEval: Towards Fast and Accurate Evaluation of Text-to-Image Diffusion Generative ModelsLin Zhao, Tianchen Zhao, Zinan Lin, Xuefei Ning 等CVPR 2024 · 被引用 2 次
