Learning Fingerprints for Medical Time Series with Redundancy-Constrained Information Maximization
Huayu Li, ZhengXiao He, Xiwen Chen, Jingjing Wang, Siyuan Tian, Jinghao Wen, Ao Li
Abstract
Learning meaningful representations from medical time series (MedTS), such as ECG or EEG signals, is a critical challenge. These signals are often high-dimensional, variable-length, and rife with noise. Existing self-supervised approaches, such as Masked Autoencoders (MAEs), are highly effective for pre-training general-purpose encoders. However, they do not explicitly learn compact, fixed-size, or semantically interpretable latent representations, typically relying on heuristic aggregation strategies such as global average pooling or a designated [CLS] token. We propose a novel framework that compresses a variable-length MedTS into a fixed-size set of latent Fingerprint Tokens. Our architecture employs a cross-attention bottleneck to generate these tokens and is trained with a dual-objective function. The first objective is a reconstruction loss, which ensures the tokens are sufficient statistics for the original data. The second, a diversity penalty based on the Total Coding Rate (TCR), explicitly minimizes the redundancy between tokens, encouraging them to become statistically disentangled representations. We present the theoretical justification for our method, framing it as a novel Disentangled Rate-Distortion problem. This approach produces a low-dimensional, interpretable, and sample-efficient representation, where each token is encouraged to capture an independent factor of variation, paving the way for more robust digital biomarkers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a79306d4-b871-41d6-94de-d45ce1730432Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
Related papers
- MedGMAE: Gaussian Masked Autoencoders for Medical Volumetric Representation LearningXueming Fu, Fenghe Tang, Rongsheng Wang, Yingtai Li et al.ICLR 2026
- Learning Cardiac Latent Representations in Vectorcardiogram SpaceBosong Huang, Panzhen Zhao, Zengxiang Li, Patricia Lee et al.ICML 2026
- Self-Supervised Dynamical System Representations for Physiological Time-SeriesYenho Chen, Maxwell A. Xu, James Rehg, Christopher RozellICML 2026 · 1 citation
- Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation ModelSam Gijsen, Marc-Andre Schulz, Kerstin RitterICLR 2026 · 5 citations
- Stochastic Optimal Control for Continuous-Time fMRI Representation LearningJoonhyeong Park, Byoungwoo Park, Chang-Bae Bang, Jungwon Choi et al.ICLR 2026 · 2 citations
