STT-LLM: Structural-Temporal Tokenization for Adapting LLMs to Longitudinal Clinical Profiles
Maxx Richard Rahman, Mostafa Hammouda, Wolfgang Maass
摘要
Large Language Models have shown strong generalization across natural language tasks but remain underexplored for longitudinal clinical profiles. In sports anti-doping, biological profiles are analyzed to support early detection of prohibited substance use and identification of anomalous biological patterns, both of which require joint modeling of temporal dynamics and metabolic relationships. We propose STT-LLM, a structural-temporal tokenization framework that adapts LLMs to longitudinal clinical analysis without modifying their backbone architectures. STT-LLM constructs biologically grounded structural-temporal embeddings and transforms them into LLM-compatible tokens via specialized tokenizers that explicitly encode pathway structure and temporal evolution. We evaluate STT-LLM on real-world longitudinal datasets from athletes, showing consistent improvements over native LLM tokenization strategies in sequence prediction and anomaly detection. In addition, we present a case study where STT-LLM provides contextual reasoning that aligns more closely with expert assessments compared to baseline models. These results highlight tokenization as a key bottleneck and opportunity for adapting LLMs to clinical data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu 等ICLR 2024 · 被引用 915 次
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
- UniTime: A Language-Empowered Unified Model for Cross-Domain Time Series ForecastingXu Liu, Junfeng Hu, Yuan Li, Shizhe Diao 等WWW 2024 · 被引用 198 次
- GPPT: Graph Pre-training and Prompt Tuning to Generalize Graph Neural NetworksMingchen Sun, Kaixiong Zhou, Xin He, Ying Wang 等KDD 2022 · 被引用 141 次
相关 Paper
- ST-LLM: Spatial Transcriptomics Embedding with Large Language ModelsZhetao Xu, Xiaohua Wan, Le Li, Shuang Feng 等AAAI 2026
- Semantic-Enhanced Time-Series Forecasting via Large Language ModelsHao Liu, Zhang xiaoxing, Chun Yang, Xiaobin ZhuICLR 2026 · 被引用 5 次
- Markovian Linguistic-Temporal Bridge: Unlocking the Potential of LLMs for Time Series ForecastingSiming Sun, Kai Zhang, Xuejun Jiang, Wenchao Meng 等ACL 2026
- Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMsKai Zhuang, Jiawei Zhang, Yumou Liu, Hanqun Cao 等ICLR 2026
- Hierarchical Graph Tokenization for Molecule-Language AlignmentYongqiang Chen, Quanming Yao, Juzheng Zhang, James Cheng 等ICML 2025 · 被引用 2 次
