Characterizing Practices, Limitations, and Opportunities Related to Text Information Extraction Workflows: A Human-in-the-loop Perspective
Sajjadur Rahman, Eser Kandogan
Abstract
Information extraction (IE) approaches often play a pivotal role in text analysis and require significant human intervention. Therefore, a deeper understanding of existing IE practices and related challenges from a human-in-the-loop perspective is warranted. In this work, we conducted semi-structured interviews in an industrial environment and analyzed the reported IE approaches and limitations. We observed that data science workers often follow an iterative task model consisting of information foraging and sensemaking loops across all the phases of an IE workflow. The task model is generalizable and captures diverse goals across these phases (e.g., data preparation, modeling, evaluation.) We found several limitations in both foraging (e.g., data exploration) and sensemaking (e.g., qualitative debugging) loops stemming from a lack of adherence to existing cognitive engineering principles. Moreover, we identified that due to the iterative nature of an IE workflow, the requirement of provenance is often implied but rarely supported by existing systems. Based on these findings, we discuss design implications for supporting IE workflows and future research directions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27f6665f-e40d-427e-b002-1c710fa96c9cBuilds on5
- Human Factors in Model Interpretability: Industry Practices, Challenges, and NeedsSungsoo Ray Hong, Jessica Hullman, Enrico BertiniCSCW 2020 · 219 citations
- Snippext: Semi-supervised Opinion Mining with Augmented DataZhengjie Miao, Yuliang Li, Xiaolan Wang, Wang-Chiew TanWWW 2020 · 94 citations
- B2: Bridging Code and Interactive Visualization in Computational NotebooksYifan Wu, Joseph M. Hellerstein, Arvind SatyanarayanUIST 2020 · 71 citations
- Parallel Worlds: Repeated Initializations of the Same Team to Improve Team ViabilityMark E. Whiting, Irena Gao, Michelle Xing, N'godjigui Junior Diarrassouba et al.CSCW 2020 · 19 citations
- NOAH: Interactive Spreadsheet Exploration with Dynamic Hierarchical OverviewsSajjadur Rahman, Mangesh Bendre, Yuyang Liu, Shichu Zhu et al.VLDB 2021 · 11 citations
Related papers
- Code Code Evolution: Understanding How People Change Data Science Notebooks Over TimeDeepthi Raghunandan, Aayushi Roy, Shenzhi Shi, Niklas Elmqvist et al.CHI 2023 · 19 citations
- "I Need to Find That One Chart": How Data Workers Navigate, Summarize and Communicate Analytical ConversationsKen Gu, Srishti Palani, Vidya SetlurCHI 2026 · 1 citation
- BugDoc: Algorithms to Debug Computational ProcessesRaoni Lourenço, Juliana Freire, Dennis E. ShashaSIGMOD 2020 · 9 citations
- Orienting, Framing, Bridging, Magic, and Counseling: How Data Scientists Navigate the Outer Loop of Client Collaborations in Industry and AcademiaSean Kross, Philip J. GuoCSCW 2021 · 35 citations
- The Influence of Visual Provenance Representations on Strategies in a Collaborative Hand-off Data Analysis ScenarioJeremy E. Block, Shaghayegh Esmaeili, Eric D. Ragan, John R. Goodall et al.IEEE VIS 2022 · 7 citations
