STT-LLM: Structural-Temporal Tokenization for Adapting LLMs to Longitudinal Clinical Profiles
Maxx Richard Rahman, Mostafa Hammouda, Wolfgang Maass
Abstract
Large Language Models have shown strong generalization across natural language tasks but remain underexplored for longitudinal clinical profiles. In sports anti-doping, biological profiles are analyzed to support early detection of prohibited substance use and identification of anomalous biological patterns, both of which require joint modeling of temporal dynamics and metabolic relationships. We propose STT-LLM, a structural-temporal tokenization framework that adapts LLMs to longitudinal clinical analysis without modifying their backbone architectures. STT-LLM constructs biologically grounded structural-temporal embeddings and transforms them into LLM-compatible tokens via specialized tokenizers that explicitly encode pathway structure and temporal evolution. We evaluate STT-LLM on real-world longitudinal datasets from athletes, showing consistent improvements over native LLM tokenization strategies in sequence prediction and anomaly detection. In addition, we present a case study where STT-LLM provides contextual reasoning that aligns more closely with expert assessments compared to baseline models. These results highlight tokenization as a key bottleneck and opportunity for adapting LLMs to clinical data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a981b5e-d94a-40f9-be04-d8458c76ddd1Builds on10
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- UniTime: A Language-Empowered Unified Model for Cross-Domain Time Series ForecastingXu Liu, Junfeng Hu, Yuan Li, Shizhe Diao et al.WWW 2024 · 198 citations
- GPPT: Graph Pre-training and Prompt Tuning to Generalize Graph Neural NetworksMingchen Sun, Kaixiong Zhou, Xin He, Ying Wang et al.KDD 2022 · 141 citations
Related papers
- ST-LLM: Spatial Transcriptomics Embedding with Large Language ModelsZhetao Xu, Xiaohua Wan, Le Li, Shuang Feng et al.AAAI 2026
- Semantic-Enhanced Time-Series Forecasting via Large Language ModelsHao Liu, Zhang xiaoxing, Chun Yang, Xiaobin ZhuICLR 2026 · 5 citations
- Markovian Linguistic-Temporal Bridge: Unlocking the Potential of LLMs for Time Series ForecastingSiming Sun, Kai Zhang, Xuejun Jiang, Wenchao Meng et al.ACL 2026
- Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMsKai Zhuang, Jiawei Zhang, Yumou Liu, Hanqun Cao et al.ICLR 2026
- Hierarchical Graph Tokenization for Molecule-Language AlignmentYongqiang Chen, Quanming Yao, Juzheng Zhang, James Cheng et al.ICML 2025 · 2 citations
