Dynamics of Human-AI Collective Knowledge on the Web: A Scalable Model and Insights for Sustainable Growth
Buddhika Nettasinghe, Kang Zhao
Abstract
Humans and large language models (LLMs) now co-produce and co-consume the web's shared knowledge archives. Such human-AI collective knowledge ecosystems contain feedback loops with both benefits (e.g., faster growth, easier learning) and systemic risks (e.g., quality dilution, skill reduction, model collapse). To understand such phenomena, we propose a minimal, interpretable dynamical model of the co-evolution of archive size, archive quality, model (LLM) skill, aggregate human skill, and query volume. The model captures two content inflows (human, LLM) controlled by a gate on LLM-content admissions, two learning pathways for humans (archive study vs. LLM assistance), and two LLM-training modalities (corpus-driven scaling vs. learning from human feedback). Through numerical experiments, we identify different growth regimes (e.g., healthy growth, inverted flow, inverted learning, oscillations), and show how platform and policy levers (gate strictness, LLM training, human learning pathways) shift the system across regime boundaries. Two domain configurations (PubMed, GitHub & Copilot) illustrate contrasting steady states under different growth rates and moderation norms. We also fit the model to Wikipedia's knowledge flow during pre-ChatGPT and post-ChatGPT eras separately. We find a rise in LLM additions with a concurrent decline in human inflow, consistent with a regime identified by the model. Our model and analysis yield actionable insights for sustainable growth of human-AI collective knowledge on the Web. CCS Concepts • Information systems → World Wide Web; • Computing methodologies → Modeling and simulation; • Human-centered computing → Collaborative and social computing theory, concepts and paradigms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f292ec6-3ff7-426a-8ca3-b0e95fc8103eBuilds on3
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Emergence of Structural Disparities in the Web of Scientific CitationsBuddhika Nettasinghe, Nazanin Alipourfard, Vikram Krishnamurthy, Kristina LermanWWW 2026 · 2 citations
- Information Retrieval in the Age of Generative AI: The RGB ModelMichele Garetto, Alessandro Cornacchia, Franco Galante, Emilio Leonardi et al.SIGIR 2025 · 2 citations
Related papers
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton et al.ICML 2024 · 123 citations
- Data Pollination: An Emergent Ecological Process Driving AI Population EvolutionShufang Xie, Qizhi Pei, Ang Lv, Jingyang Hu et al.ACL 2026
- The Lock-in Hypothesis: Stagnation by AlgorithmTianyi Qiu, Zhonghao He, Tejasveer Chugh, Max Kleiman-WeinerICML 2025
- LLMs in Wikipedia: Investigating How LLMs Impact Participation in Knowledge CommunitiesMoyan Zhou, Soobin Cho, Loren TerveenCSCW 2026
- Are LLM Web Search Engines Sustainable? A Web-Measurement Study of Real-Time FetchingAbdur-Rahman Ibrahim Sayyid-Ali, Daanish Uddin Khan, Naveed Anwar BhattiWWW 2026
