Make It Count: Text-to-Image Generation with an Accurate Number of Objects
Lital Binyamin, Yoad Tewel, Hilit Segev, Eran Hirsch, Royi Rassin, Gal Chechik
2025年份
17顶会引用
摘要
CountGen (ours) "A photo of six kittens sitting on a branch" "A photo of five eggs in a carton" "A realistic photo of Goldilocks and three bears eating a porridge" "an illustration of four ninja turtles" SDXL "A realistic photo of seven dwarves dancing in the forest" Figure 1 . CountGen generates the correct number of objects specified in the input prompt while maintaining a natural layout that aligns with the prompt.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video GenerationAriel Shaulov, Itay Hazan, Lior Wolf, Hila CheferNeurIPS 2025 · 被引用 22 次
- SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image GenerationSashuai zhou, Qiang Zhou, Ma Junpeng, Yue Cao 等CVPR 2026 · 被引用 7 次
- Boosting Quantitive and Spatial Awareness for Zero-Shot Object CountingDa Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao 等CVPR 2026 · 被引用 6 次
- YOLO-Count: Differentiable Object Counting for Text-to-Image GenerationGuanning Zeng, Xiang Zhang, Zirui Wang, Haiyang Xu 等ICCV 2025 · 被引用 4 次
- Be Decisive: Noise-Induced Layouts for Multi-Subject GenerationOmer Dahary, Yehonathan Cohen, Or Patashnik, Kfir Aberman 等SIGGRAPH 2025 · 被引用 3 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani 等NeurIPS 2023 · 被引用 462 次
- Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion ModelsHila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf 等SIGGRAPH 2023 · 被引用 438 次
相关 Paper
- TextCraftor: Your Text Encoder can be Image Quality ControllerYanyu Li, Xian Liu, Anil Kag, Ju Hu 等CVPR 2024
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
- DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven GenerationNataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch 等CVPR 2023
- Check, Locate, Rectify: A Training-Free Layout Calibration System for Text- to- Image GenerationBiao Gong, Siteng Huang, Yutong Feng, Shiwei Zhang 等CVPR 2024
- CountGD++: Generalized Prompting for Open-World CountingNiki Amini-Naieni, Andrew ZissermanCVPR 2026 · 被引用 14 次
