HICode: Hierarchical Inductive Coding with LLMs
Mian Zhong, Pristina Wang, Anjalie Field
摘要
Despite numerous applications for fine-grained corpus analysis, researchers continue to rely on manual labeling, which does not scale, or statistical tools like topic modeling, which are difficult to control. We propose that LLMs have the potential to scale the nuanced analyses that researchers typically conduct manually to large text corpora. To this effect, inspired by qualitative research methods, we develop HICode 1 , a two-part pipeline that first inductively generates labels directly from analysis data and then hierarchically clusters them to surface emergent themes. We validate this approach across three diverse datasets by measuring alignment with human-constructed themes and demonstrating its robustness through automated and human evaluations. Finally, we conduct a case study of litigation documents related to the ongoing opioid crisis in the U.S., revealing aggressive marketing strategies employed by pharmaceutical companies and demonstrating HICode's potential for facilitating nuanced analyses in large-scale data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceAlexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov 等NeurIPS 2021 · 被引用 220 次
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 被引用 173 次
- Using Thematic Analysis in Healthcare HCI at CHI: A Scoping ReviewRobert Bowman, Camille Nadal, Kellie Morrissey, Anja Thieme 等CHI 2023 · 被引用 106 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language ModelsJie Gao, Yuchen Guo, Gionnieve Lim, Tianqin Zhang 等CHI 2024 · 被引用 63 次
相关 Paper
- Scholastic: Graphical Human-AI Collaboration for Inductive and Interpretive Text AnalysisMatt-Heun Hong, Lauren A. Marsh, Jessica L. Feuston, Janet Ruppert 等UIST 2022 · 被引用 26 次
- MedLinkDE - MedDRA Entity Linking for German with Guided Chain of Thought ReasoningRoman Christof, Farnaz Zeidi, Manuela Messelhäußer, Dirk Mentzer 等EMNLP 2025
- End-To-End Causal Effect Estimation from Unstructured Natural Language DataNikita Dhawan, Leonardo Cotta, Karen Ullrich, Rahul G. Krishnan 等NeurIPS 2024 · 被引用 24 次
- Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic ModelsZongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu 等ACL 2025
- Thematic-LM: A LLM-based Multi-agent System for Large-scale Thematic AnalysisTingrui Qiao, Caroline Walker, Chris Cunningham, Yun Sing KohWWW 2025 · 被引用 32 次
