PPGPT: Transferring Next-Token Modeling from Language to PPG Signals
Zexing Zhang, Huimin Lu, Qingxin Zhao
Abstract
The success of large language models (LLMs) in cognitive tasks prompts the question of whether their next-token prediction (NTP) paradigm can be adapted to model physiological signals from wearable devices. A key target for this adaptation is photoplethysmography (PPG), the most prevalent sensing modality in consumer wearables for non-invasive monitoring of diverse physiological conditions. Unlike in NLP, where NTP aligns with generative objectives, physiological signal analysis involves fundamentally different tasks, such as continuous parameter estimation (regression) and discrete state recognition (classification). This disparity creates a semantic mismatch between the pre-training paradigm and the downstream tasks. To bridge this gap, we propose PPGPT, the first foundation model that reformulates NTP into next-feature token prediction (NFTP), learning hierarchical feature transition probabilities to unify pre-training and downstream objectives. PPGPT features a novel dual-stream encoder that generates feature tokens by jointly modeling temporal dynamics and local-global morphological patterns. The model is developed using a two-stage training framework: it is first pre-trained on a large-scale mixed dataset of 1.6 billion data points and then validated on our newly released BioMTL benchmark, which includes data from 172 subjects over 285 days across seven different tasks. Extensive experiments show that PPGPT significantly outperforms competing methods, achieving a 16.5% improvement in F1-score and a 25.9% reduction in Mean Absolute Error (MAE). Furthermore, the model demonstrates robust few-shot learning capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7ce66fd-a28e-411a-8376-720c8331e97aBuilds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIWei-Bang Jiang, Li-Ming Zhao, Bao-Liang LuICLR 2024 · 298 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
- RankMe: Assessing the Downstream Performance of Pretrained Self-Supervised Representations by Their RankQuentin Garrido, Randall Balestriero, Laurent Najman, Yann LeCunICML 2023 · 127 citations
Related papers
- A robust PPG foundation model using multimodal physiological supervisionEloy Geenjaar, Vince Calhoun, scott daly, Gouthaman KV et al.ICML 2026 · 1 citation
- PaPaGei: Open Foundation Models for Optical Physiological SignalsArvind Pillai, Dimitris Spathis, Fahim Kawsar, Mohammad MalekzadehICLR 2025
- Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field SettingsMithun Saha, Maxwell A. Xu, Wanting Mao, Sameer Neupane et al.UbiComp 2025 · 17 citations
- PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological SensingYiping Xie, Bo Zhao, Mingtong Dai, Jian-Ping Zhou et al.ICLR 2026 · 19 citations
- Beyond Hearing: Learning Task-Agnostic ExG Representations from Earphones via Physiology-Informed TokenizationHyungjun Yoon, Seungjoo Lee, Yu Wu, Xiaomeng Chen et al.ICLR 2026
