Learning Disentangled Representation for Multi-Modal Time-Series Sensing Signals
Ruichu Cai, Zhifan Jiang, Kaitao Zheng, Zijian Li, Weilin Chen, Xuexin Chen, Yifan Shen, Guangyi Chen, Zhifeng Hao, Kun Zhang
摘要
Existing methods for multi-modal time series representation learning aim to disentangle the modality-shared and modality-specific latent variables. Although achieving notable performances on downstream tasks, they usually assume an orthogonal latent space. However, the modality-specific and modality-shared latent variables might be dependent on real-world scenarios. Therefore, we propose a general generation process, where the modality-shared and modality-specific latent variables are dependent, and further develop a Multi-modAl TEmporal Disentanglement (MATE) model. Specifically, our MATE model is built on a temporally variational inference architecture with the modality-shared and modality-specific prior networks for the disentanglement of latent variables. Furthermore, we establish identifiability results to show that the extracted representation is disentangled. More specifically, we first achieve the subspace identifiability for modality-shared and modality-specific latent variables by leveraging the pairing of multi-modal data. Then we establish the component-wise identifiability of modality-specific latent variables by employing sufficient changes of historical latent variables. Extensive experimental studies on multi-modal sensors, human activity recognition, and healthcare datasets show a general improvement in different downstream tasks, highlighting the effectiveness of our method in real-world scenarios. * Equal contributions Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
- TS2Vec: Towards Universal Representation of Time SeriesZhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang 等AAAI 2022 · 被引用 938 次
相关 Paper
- On Hierarchical Disentanglement of Interactive Behaviors for Multimodal Spatiotemporal Data with IncompletenessJiayi Chen, Aidong ZhangKDD 2023 · 被引用 4 次
- Learning Disentangled Factors from Paired Data in Cross-Modal Retrieval: An Implicit Identifiable VAE ApproachMinyoung Kim, Ricardo Guerrero, Vladimir PavlovicACM MM 2021 · 被引用 4 次
- Disentangling Multi-view Representations Beyond Inductive BiasGuanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang 等ACM MM 2023 · 被引用 15 次
- Scaling Up Multivariate Time Series Pre-Training with Decoupled Spatial-Temporal RepresentationsRui Zha, Le Zhang, Shuangli Li, Jingbo Zhou 等ICDE 2024 · 被引用 9 次
- Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal RetrievalJialin Tian, Kai Wang, Xing Xu, Zuo Cao 等SIGIR 2022 · 被引用 19 次
