Representative Language Generation
Charlotte Peale, Vinod Raman, Omer Reingold
摘要
We introduce "representative generation," extending the theoretical framework for generation proposed by Kleinberg et al. (2024) and formalized by Li et al. (2024) , to additionally address diversity and bias concerns in generative models. Our notion requires outputs of a generative model to proportionally represent groups of interest from the training data. We characterize representative uniform and non-uniform generation, introducing the "group closure dimension" as a key combinatorial quantity. For representative generation in the limit, we analyze both information-theoretic and computational aspects, demonstrating feasibility for countably infinite hypothesis classes and collections of groups under certain conditions, but proving a negative result for computability using only membership queries. This contrasts with Kleinberg et al.'s (2024) positive results for standard generation in the limit. Our findings provide a rigorous foundation for developing more diverse and representative generative models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- On Union-Closedness of Language GenerationSteve Hanneke, Amin Karbasi, Anay Mehrotra, Grigoris VelegkasNeurIPS 2025 · 被引用 17 次
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 被引用 4 次
- Characterizing the Effect of Noise in Language Generation in the LimitAaron Li, Ian ZhangICML 2026 · 被引用 4 次
- Language Generation in the Limit: Noise, Loss, and FeedbackYannan Bai, Debmalya Panigrahi, Ian ZhangSODA 2026
- Language Identification in the Limit with Computational TraceBinghui Peng, Amin Saberi, Grigoris VelegkasICLR 2026
它引用的顶会 Paper7
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language ModelsHannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal 等NeurIPS 2021 · 被引用 243 次
- Language Generation in the LimitJon M. Kleinberg, Sendhil MullainathanNeurIPS 2024 · 被引用 45 次
- Outcome indistinguishabilityCynthia Dwork, Michael P. Kim, Omer Reingold, Guy N. Rothblum 等STOC 2021 · 被引用 24 次
- Taming Mode Collapse in Score Distillation for Text-to-3D GenerationPeihao Wang, Dejia Xu, Zhiwen Fan, Dilin Wang 等CVPR 2024 · 被引用 8 次
- On the Limits of Language Generation: Trade-Offs between Hallucination and Mode-CollapseAlkis Kalavasis, Anay Mehrotra, Grigoris VelegkasSTOC 2025 · 被引用 2 次
相关 Paper
- Multi-Group Proportional Representations for Text-to-Image ModelsSangwon Jung, Alex Oesterling, Claudio Mayrink Verdun, Sajani Vithana 等CVPR 2025
- A Fair Generative Model Using LeCam DivergenceSoobin Um, Changho SuhAAAI 2023 · 被引用 8 次
- Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-VotingPreethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi 等EMNLP 2023 · 被引用 14 次
- Generation from Noisy ExamplesAnanth Raman, Vinod RamanICML 2025
- Group-Aware Reinforcement Learning for Output Diversity in Large Language ModelsOron Anschel, Alon Shoshan, Adam Botach, Shunit Haviv Hakimi 等EMNLP 2025 · 被引用 1 次
