Learning Fingerprints for Medical Time Series with Redundancy-Constrained Information Maximization
Huayu Li, ZhengXiao He, Xiwen Chen, Jingjing Wang, Siyuan Tian, Jinghao Wen, Ao Li
摘要
Learning meaningful representations from medical time series (MedTS), such as ECG or EEG signals, is a critical challenge. These signals are often high-dimensional, variable-length, and rife with noise. Existing self-supervised approaches, such as Masked Autoencoders (MAEs), are highly effective for pre-training general-purpose encoders. However, they do not explicitly learn compact, fixed-size, or semantically interpretable latent representations, typically relying on heuristic aggregation strategies such as global average pooling or a designated [CLS] token. We propose a novel framework that compresses a variable-length MedTS into a fixed-size set of latent Fingerprint Tokens. Our architecture employs a cross-attention bottleneck to generate these tokens and is trained with a dual-objective function. The first objective is a reconstruction loss, which ensures the tokens are sufficient statistics for the original data. The second, a diversity penalty based on the Total Coding Rate (TCR), explicitly minimizes the redundancy between tokens, encouraging them to become statistically disentangled representations. We present the theoretical justification for our method, framing it as a novel Disentangled Rate-Distortion problem. This approach produces a low-dimensional, interpretable, and sample-efficient representation, where each token is encouraged to capture an independent factor of variation, paving the way for more robust digital biomarkers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
相关 Paper
- MedGMAE: Gaussian Masked Autoencoders for Medical Volumetric Representation LearningXueming Fu, Fenghe Tang, Rongsheng Wang, Yingtai Li 等ICLR 2026
- Learning Cardiac Latent Representations in Vectorcardiogram SpaceBosong Huang, Panzhen Zhao, Zengxiang Li, Patricia Lee 等ICML 2026
- Self-Supervised Dynamical System Representations for Physiological Time-SeriesYenho Chen, Maxwell A. Xu, James Rehg, Christopher RozellICML 2026 · 被引用 1 次
- Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation ModelSam Gijsen, Marc-Andre Schulz, Kerstin RitterICLR 2026 · 被引用 5 次
- Stochastic Optimal Control for Continuous-Time fMRI Representation LearningJoonhyeong Park, Byoungwoo Park, Chang-Bae Bang, Jungwon Choi 等ICLR 2026 · 被引用 2 次
