From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP
Adithya V. Ganesan, Vasudha Varadarajan, Oscar N. E. Kjell, Whitney Ringwald, Scott M. Feltman, Benjamin J. Luft, Roman Kotov, Ryan L. Boyd, H. Andrew Schwartz
Abstract
While NLP typically treats documents as independent and unordered samples, in longitudinal studies, this assumption rarely holds: documents are nested within authors and ordered in time, forming person-indexed, time-ordered . Here, we demonstrate the need for and propose a longitudinal modeling and evaluation paradigm that consequently updates four parts of the NLP pipeline: (1) evaluation splits aligned to generalization over people () and/or time (); (2) accuracy metrics separating between-person differences from within-person dynamics; (3) sequence inputs to incorporate history by default; and (4) model internals that support different of latent state over histories (pooled summaries, explicit dynamics, or interaction-based models). We demonstrate the issues ensued by traditional pipeline and our proposed improvements on a dataset of 17k daily diary transcripts paired with PTSD symptom severity from 238 participants, finding that traditional document-level evaluation can yield substantially different and sometimes reversed conclusions compared to our ecologically valid modeling and evaluation. We tie our results to a broader discussion motivating a shift from word-sequence evaluation toward paradigms for NLP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb7f296b-d29c-4be6-8c55-9ce22e4f77aeBuilds on9
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Quantifying Privacy Risks of Masked Language Models Using Membership Inference AttacksFatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick et al.EMNLP 2022 · 72 citations
- Identifying Moments of Change from Longitudinal User TextAdam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim et al.ACL 2022 · 46 citations
- Leveraging Similar Users for Personalized Language Modeling with Limited DataCharles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas et al.ACL 2022 · 37 citations
Related papers
- Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length ContextsYuho Lee, Jiaqi Deng, Nicole Hee-Yeon Kim, Hyangsuk Min et al.EMNLP 2025
- LENS: LLM-Enabled Narrative Synthesis for Mental Health by Aligning Multimodal Sensing with Language ModelsWenxuan Xu, Arvind Pillai, Subigya Nepal, Amanda C. Collins et al.ACL 2026 · 1 citation
- Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular DataLucas Rosenblatt, Peihan Liu, Ryan McKenna, Natalia PonomarevaICML 2026 · 1 citation
- Temporal reasoning for timeline summarisation in social mediaJiayu Song, Mahmud Elahi Akhter, Dana Atzil-Slonim, Maria LiakataACL 2025 · 6 citations
- MIND: Empowering Mental Health Clinicians with Multimodal Data Insights through a Narrative DashboardRuishi Zou, Shiyu Xu, Margaret E. Morris, Jihan Ryu et al.CHI 2026 · 2 citations
