Existence and Estimation of Critical Batch Size for Training Generative Adversarial Networks with Two Time-Scale Update Rule
Naoki Sato, Hideaki Iiduka
摘要
Previous results have shown that a two time-scale update rule (TTUR) using different learning rates, such as different constant rates or different decaying rates, is useful for training generative adversarial networks (GANs) in theory and in practice. Moreover, not only the learning rate but also the batch size is important for training GANs with TTURs and they both affect the number of steps needed for training. This paper studies the relationship between batch size and the number of steps needed for training GANs with TTURs based on constant learning rates. We theoretically show that, for a TTUR with constant learning rates, the number of steps needed to find stationary points of the loss functions of both the discriminator and generator decreases as the batch size increases and that there exists a critical batch size minimizing the stochastic first-order oracle (SFO) complexity. Then, we use the Fr'echet inception distance (FID) as the performance measure for training and provide numerical results indicating that the number of steps needed to achieve a low FID score decreases as the batch size increases and that the SFO complexity increases once the batch size exceeds the measured critical batch size. Moreover, we show that measured critical batch sizes are close to the sizes estimated from our theoretical results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda 等NeurIPS 2020 · 被引用 697 次
- Projected GANs Converge FasterAxel Sauer, Kashyap Chitta, Jens Müller, Andreas GeigerNeurIPS 2021 · 被引用 325 次
- On Aliased Resizing and Surprising Subtleties in GAN EvaluationGaurav Parmar, Richard Zhang, Jun-Yan ZhuCVPR 2022 · 被引用 250 次
- COT-GAN: Generating Sequential Data via Causal Optimal TransportTianlin Xu, Li Kevin Wenliang, Michael Munn, Beatrice AcciaioNeurIPS 2020 · 被引用 139 次
- Low-Rank Subspaces in GANsJiapeng Zhu, Ruili Feng, Yujun Shen, Deli Zhao 等NeurIPS 2021 · 被引用 80 次
相关 Paper
- Online Learning and Information Exponents: The Importance of Batch size & Time/Complexity TradeoffsLuca Arnaboldi, Yatin Dandi, Florent Krzakala, Bruno Loureiro 等ICML 2024 · 被引用 4 次
- ACT-Diffusion: Efficient Adversarial Consistency Training for One-Step Diffusion ModelsFei Kong, Jinhao Duan, Lichao Sun, Hao Cheng 等CVPR 2024
- From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGDKonstantinos C. Tsiolis, Alireza Mousavi-Hosseini, Murat A. ErdogduNeurIPS 2025 · 被引用 2 次
- Top-k Training of GANs: Improving GAN Performance by Throwing Away Bad SamplesSamarth Sinha, Zhengli Zhao, Anirudh Goyal, Colin Raffel 等NeurIPS 2020 · 被引用 48 次
- Adaptive Weighted Discriminator for Training Generative Adversarial NetworksVasily Zadorozhnyy, Qiang Cheng, Qiang YeCVPR 2021
