Collaboration or Corporate Capture? Quantifying NLP's Reliance on Industry Artifacts and Contributions
Will Aitken, Mohamed Abdalla, Karen Rudie, Catherine Stinson
Abstract
Impressive performance of pre-trained models has garnered public attention and made news headlines in recent years. Almost always, these models are produced by or in collaboration with industry. Using them is critical for competing on natural language processing (NLP) benchmarks and correspondingly to stay relevant in NLP research. We surveyed 100 papers published at EMNLP 2022 to determine the degree to which researchers rely on industry models, other artifacts, and contributions to publish in prestigious NLP venues and found that the ratio of their citation is at least three times greater than what would be expected. Our work serves as a scaffold to enable future researchers to more accurately address whether: 1) Collaboration with industry is still collaboration in the absence of an alternative or 2) if NLP inquiry has been captured by the motivations and research direction of private corporations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing ResearchMohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aurélie Névéol et al.ACL 2023 · 16 citations
- What Do NLP Researchers Believe? Results of the NLP Community MetasurveyJulian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller et al.ACL 2023 · 16 citations
- To Build Our Future, We Must Know Our Past: Contextualizing Paradigm Shifts in Natural Language ProcessingSireesh Gururaja, Amanda Bertsch, Clara Na, David Gray Widder et al.EMNLP 2023 · 6 citations
- We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic FieldsJan Philip Wahle, Terry Ruas, Mohamed Abdalla, Bela Gipp et al.EMNLP 2023 · 6 citations
Related papers
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp et al.ICML 2024 · 213 citations
- Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering StudiesFlorian Angermeir, Maximilian Amougou, Mark Kreitz, Andreas Bauer et al.ICSE 2026 · 1 citation
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 18 citations
- Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and ConsiderationsKarla Felix Navarro, Eugene Syriani, Ian ArawjoCHI 2026 · 1 citation
- What Do Position Embeddings Learn? An Empirical Study of Pre-Trained Language Model Positional EncodingYu-An Wang, Yun-Nung ChenEMNLP 2020 · 71 citations
