CountsDiff: A diffusion model on the natural numbers for generation and imputation of count-based data
Renzo Soatto, Anders Hoel, Greycen Ren, Shorna Alam, Stephen Bates, Nikolaos Daskalakis, Caroline Uhler, Maria Skoularidou
摘要
Diffusion models have excelled at generative tasks for both continuous and token-based domains, but their application to discrete ordinal data remains underdeveloped. We present CountsDiff, a diffusion framework designed to model distributions on the natural numbers. CountsDiff extends the Blackout diffusion framework by simplifying its formulation through a direct parameterization in terms of a survival probability schedule and an explicit loss weighting. This introduces flexibility through design parameters with direct analogues in existing diffusion modeling frameworks. Beyond this reparameterization, CountsDiff introduces features from modern diffusion models, previously absent in counts-based domains, including continuous-time training, classifier-free guidance, and churn/remasking reverse dynamics that allow non-monotone reverse trajectories. We propose an initial instantiation of CountsDiff and validate it on natural image datasets (CIFAR-10, CelebA), exploring the effects of the introduced design parameters in a complex, well-studied, and interpretable data domain. We then highlight biological count assays as a natural use case, evaluating CountsDiff on single-cell RNA-seq imputation in fetal and heart cell atlases. Remarkably, we find that even this simple instantiation matches or surpasses the performance of a state-of-the-art discrete generative model and leading scRNA-seq imputation methods, while leaving substantial headroom for further gains through optimized design choices in future work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and ExpressionMingxuan Wang, Gaoyang Jiang, ZiJia Ren, Cheng Chen 等ICML 2026
- Blackout Diffusion: Generative Diffusion Models in Discrete-State SpacesJavier E. Santos, Zachary R. Fox, Nicholas Lubbers, Yen Ting LinICML 2023 · 被引用 28 次
- Count Bridges enable Modeling and Deconvolving Transcriptomic DataNic Fishman, Gokul Gowri, Tanush Kumar, Jiaqi Lu 等ICLR 2026 · 被引用 3 次
- Scalable Single-Cell Gene Expression Generation with Latent Diffusion ModelsGiovanni Palla, Sudarshan Babu, Payam Dibaeinia, James Pearce 等ICML 2026 · 被引用 7 次
- Discrete Modeling via Boundary Conditional Diffusion ProcessesYuxuan Gu, Xiaocheng Feng, Lei Huang, Yingsheng Wu 等NeurIPS 2024
