HICode: Hierarchical Inductive Coding with LLMs
Mian Zhong, Pristina Wang, Anjalie Field
Abstract
Despite numerous applications for fine-grained corpus analysis, researchers continue to rely on manual labeling, which does not scale, or statistical tools like topic modeling, which are difficult to control. We propose that LLMs have the potential to scale the nuanced analyses that researchers typically conduct manually to large text corpora. To this effect, inspired by qualitative research methods, we develop HICode 1 , a two-part pipeline that first inductively generates labels directly from analysis data and then hierarchically clusters them to surface emergent themes. We validate this approach across three diverse datasets by measuring alignment with human-constructed themes and demonstrating its robustness through automated and human evaluations. Finally, we conduct a case study of litigation documents related to the ongoing opioid crisis in the U.S., revealing aggressive marketing strategies employed by pharmaceutical companies and demonstrating HICode's potential for facilitating nuanced analyses in large-scale data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceAlexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov et al.NeurIPS 2021 · 220 citations
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 173 citations
- Using Thematic Analysis in Healthcare HCI at CHI: A Scoping ReviewRobert Bowman, Camille Nadal, Kellie Morrissey, Anja Thieme et al.CHI 2023 · 106 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language ModelsJie Gao, Yuchen Guo, Gionnieve Lim, Tianqin Zhang et al.CHI 2024 · 63 citations
Related papers
- Scholastic: Graphical Human-AI Collaboration for Inductive and Interpretive Text AnalysisMatt-Heun Hong, Lauren A. Marsh, Jessica L. Feuston, Janet Ruppert et al.UIST 2022 · 26 citations
- MedLinkDE - MedDRA Entity Linking for German with Guided Chain of Thought ReasoningRoman Christof, Farnaz Zeidi, Manuela Messelhäußer, Dirk Mentzer et al.EMNLP 2025
- End-To-End Causal Effect Estimation from Unstructured Natural Language DataNikita Dhawan, Leonardo Cotta, Karen Ullrich, Rahul G. Krishnan et al.NeurIPS 2024 · 24 citations
- Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic ModelsZongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu et al.ACL 2025
- Thematic-LM: A LLM-based Multi-agent System for Large-scale Thematic AnalysisTingrui Qiao, Caroline Walker, Chris Cunningham, Yun Sing KohWWW 2025 · 32 citations
