YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
Guanning Zeng, Xiang Zhang, Zirui Wang, Haiyang Xu, Zeyuan Chen, Bingnan Li, Zhuowen Tu
Abstract
We propose YOLO-Count, a differentiable open-vocabulary object counting model that tackles both general counting challenges and enables precise quantity control for text-to-image (T2I) generation. A core contribution is the 'cardinality' map, a novel regression target that accounts for variations in object size and spatial distribution. Leveraging representation alignment and a hybrid strong-weak supervision scheme, YOLO-Count bridges the gap between open-vocabulary counting and T2I generation control. Its fully differentiable architecture facilitates gradient-based optimization, enabling accurate object count estimation and fine-grained guidance for generative models. Extensive experiments demonstrate that YOLO-Count achieves state-of-the-art counting accuracy while providing robust and effective quantity control for T2I systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88301999-e982-4659-bab1-bfb0ad537582Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsShihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao et al.NeurIPS 2023 · 505 citations
Related papers
- CountGD: Multi-Modal Open-World CountingNiki Amini-Naieni, Tengda Han, Andrew ZissermanNeurIPS 2024 · 96 citations
- T2ICount: Enhancing Cross-modal Understanding for Zero-Shot CountingYifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei et al.CVPR 2025
- CLIP-Count: Towards Text-Guided Zero-Shot Object CountingRuixiang Jiang, Lingbo Liu, Changwen ChenACM MM 2023 · 78 citations
- Improving Controllable Generation: Faster Training and Better Performance via x0-SupervisionAmadou S. Sangare, Adrien Maglo, Mohamed Chaouch, Bertrand LuvisonCVPR 2026 · 2 citations
- InstanceGen: Image Generation with Instance-level InstructionsEtai Sella, Yanir Kleiman, Hadar Averbuch-ElorSIGGRAPH 2025 · 2 citations
