Where Do People Tell Stories Online? Story Detection Across Online Communities
Maria Antoniak, Joel Mire, Maarten Sap, Elliott Ash, Andrew Piper
Abstract
Story detection in online communities is a challenging task as stories are scattered across communities and interwoven with non-storytelling spans within a single text. We address this challenge by building and releasing the StorySeeker toolkit, including a richly annotated dataset of 502 Reddit posts and comments, a detailed codebook adapted to the social media context, and models to predict storytelling at the document and span levels. Our dataset is sampled from hundreds of popular Englishlanguage Reddit communities ranging across 33 topic categories, and it contains fine-grained expert annotations, including binary story labels, story spans, and event spans. We evaluate a range of detection methods using our data, and we identify the distinctive textual features of online storytelling, focusing on storytelling spans. We illuminate distributional characteristics of storytelling on a large communitycentric social media platform, and we also conduct a case study on r/ChangeMyView, where storytelling is used as one of many persuasive strategies, illustrating that our data and models can be used for both inter-and intra-community research. Finally, we discuss implications of our tools and analyses for narratology and the study of online communities. * introductory text about the subreddit, why they're posting, etc. * questions about the story * explanations, discussion, hypotheses external to the story -Ask yourself: Is this text necessary if I were writing a summary of the story? • When highlighting the event spans:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fc254f2-a416-464f-ae2b-ae72b4bd1387Cited by top-tier papers5
- AUTALIC: A Dataset for Anti-AUTistic Ableist Language In ContextNaba Rizvi, Harper Strickland, Daniel Gitelman, Alexis Morales Flores et al.ACL 2025 · 5 citations
- Social Story Frames: Contextual Reasoning about Narrative Intent and ReceptionJoel Mire, Maria Antoniak, Steven R. Wilson, Zexin Ma et al.ACL 2026 · 2 citations
- Confabulation: The Surprising Value of Large Language Model HallucinationsPeiqi Sui, Eamon Duede, Sophie Wu, Richard Jean SoACL 2024
- Fora: A corpus and framework for the study of facilitated dialogueHope Schroeder, Deb Roy, Jad KabbaraACL 2024
- A Structured Clustering Approach for Inducing Media NarrativesRohan Das, Advait Deshmukh, Alexandria Leto, Zohar Naaman et al.ACL 2026
Builds on4
- Narrative Theory for Computational Narrative UnderstandingAndrew Piper, Richard Jean So, David BammanEMNLP 2021 · 62 citations
- Dataset Cartography: Mapping and Diagnosing Datasets with Training DynamicsSwabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang et al.EMNLP 2020 · 12 citations
- StoryARG: a corpus of narratives and personal experiences in argumentative textsNeele Falk, Gabriella LapesaACL 2023 · 1 citation
- Reports of personal experiences and stories in argumentation: datasets and analysisNeele Falk, Gabriella LapesaACL 2022
Related papers
- The Empirical Variability of Narrative Perceptions of Social Media TextsJoel Mire, Maria Antoniak, Elliott Ash, Andrew Piper et al.EMNLP 2024 · 1 citation
- STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story GenerationNader Akoury, Shufan Wang, Josh Whiting, Stephen Hood et al.EMNLP 2020 · 5 citations
- STORYSUMM: Evaluating Faithfulness in Story SummarizationMelanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams et al.EMNLP 2024 · 2 citations
- PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News ArticlesNikolaos Nikolaidis, Nicolas Stefanovitch, Purificação Silvano, Dimitar Iliyanov Dimitrov et al.ACL 2025
- Classifying Unreliable Narrators with Large Language ModelsAnneliese Brei, Katharine Henry, Abhisheik Sharma, Shashank Srivastava et al.ACL 2025
