Learning Mixtures of Markov Chains and MDPs
Chinmaya Kausik, Kevin Tan, Ambuj Tewari
摘要
We present an algorithm for learning mixtures of Markov chains and Markov decision processes (MDPs) from short unlabeled trajectories. Specifically, our method handles mixtures of Markov chains with optional control input by going through a multi-step process, involving (1) a subspace estimation step, (2) spectral clustering of trajectories using "pairwise distance estimators," along with refinement using the EM algorithm, (3) a model estimation step, and (4) a classification step for predicting labels of new trajectories. We provide end-to-end performance guarantees, where we only explicitly require the length of trajectories to be linear in the number of states and the number of trajectories to be linear in a mixing time parameter. Experimental results support these guarantees, where we attain 96.6% average accuracy on a mixture of two MDPs in gridworld, outperforming the EM algorithm with random initialization (73.2% average accuracy).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement LearningShuguang Yu, Shuxing Fang, Ruixin Peng, Zhengling Qi 等NeurIPS 2024 · 被引用 9 次
- RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy EvaluationJeongyeol Kwon, Shie Mannor, Constantine Caramanis, Yonathan EfroniNeurIPS 2024 · 被引用 9 次
- Learning Mixtures of Markov Chains with Quality GuaranteesFabian Spaeh, Charalampos E. TsourakakisWWW 2023 · 被引用 5 次
- Test-Time Regret Minimization in Meta Reinforcement LearningMirco Mutti, Aviv TamarICML 2024 · 被引用 4 次
- Markovletics: Methods and A Novel Application for Learning Continuous-Time Markov Chain MixturesFabian Spaeh, Charalampos E. TsourakakisWWW 2024 · 被引用 2 次
它引用的顶会 Paper5
- RL for Latent MDPs: Regret Guarantees and a Lower BoundJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 被引用 91 次
- Meta-learning for Mixed Linear RegressionWeihao Kong, Raghav Somani, Zhao Song, Sham M. Kakade 等ICML 2020 · 被引用 70 次
- Learning Mixtures of Linear Dynamical SystemsYanxi Chen, H. Vincent PoorICML 2022 · 被引用 22 次
- Coresets for Time Series ClusteringLingxiao Huang, K. Sudhir, Nisheeth K. VishnoiNeurIPS 2021 · 被引用 22 次
- Model-Free and Model-Based Policy Evaluation when Causality is UncertainDavid Bruns-SmithICML 2021 · 被引用 14 次
相关 Paper
- Spectral Learning for Infinite-Horizon Average-Reward POMDPsAlessio Russo, Alberto Maria Metelli, Marcello RestelliNeurIPS 2025
- Deep Dirichlet Process Mixture Model for Non-parametric Trajectory ClusteringDi Yao, Jin Wang, Wenjie Chen, Fangda Guo 等ICDE 2024 · 被引用 7 次
- Fast Mixing Steady-State Control in Markov Decision ProcessesFederico Corso, Marco Mussi, Alberto Maria MetelliICML 2026
- Regime Switching BanditsXiang Zhou, Yi Xiong, Ningyuan Chen, Xuefeng GaoNeurIPS 2021 · 被引用 23 次
- Impact of Connectivity on Laplacian Representations in Reinforcement LearningTommaso Giorgi, Pierriccardo Olivieri, Keyue Jiang, Laura Toni 等ICML 2026 · 被引用 1 次
