Pre-trained Large Language Models Learn to Predict Hidden Markov Models In-context
Yijia Dai, Zhaolin Gao, Yahya Sattar, Sarah Dean, Jennifer J. Sun
摘要
Hidden Markov Models (HMMs) are foundational tools for modeling sequential data with latent Markovian structure, yet fitting them to real-world data remains computationally challenging. In this work, we show that pre-trained large language models (LLMs) can effectively model data generated by HMMs via in-context learning (ICL)-their ability to infer patterns from examples within a prompt. On a diverse set of synthetic HMMs, LLMs achieve predictive accuracy approaching the theoretical optimum. We uncover novel scaling trends influenced by HMM properties, and offer theoretical conjectures for these empirical observations. We also provide practical guidelines for scientists on using ICL as a diagnostic tool for complex data. On real-world animal decision-making tasks, ICL achieves competitive performance with models designed by human experts. To our knowledge, this is the first demonstration that ICL can learn to predict HMM-generated sequences-an advance that deepens our understanding of in-context learning in LLMs and establishes its potential as a powerful tool for uncovering hidden structure in complex scientific data. Our code is available at https://github.com/DaiYijia02/icl-hmm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang 等NeurIPS 2022 · 被引用 407 次
- Are Language Models Actually Useful for Time Series Forecasting?Mingtian Tan, Mike A. Merrill, Vinayak Gupta, Tim Althoff 等NeurIPS 2024 · 被引用 326 次
相关 Paper
- Task Descriptors Help Transformers Learn Linear Models In-ContextRuomin Huang, Rong GeICLR 2025
- What Do Language Models Learn in Context? The Structured Task HypothesisJiaoda Li, Yifan Hou, Mrinmaya Sachan, Ryan CotterellACL 2024 · 被引用 5 次
- Compositional Exemplars for In-context LearningJiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu 等ICML 2023 · 被引用 188 次
- Parallel Structures in Pre-training Data Yield In-Context LearningYanda Chen, Chen Zhao, Zhou Yu, Kathleen R. McKeown 等ACL 2024 · 被引用 1 次
- In-Context Learning Learns Label Relationships but Is Not Conventional LearningJannik Kossen, Yarin Gal, Tom RainforthICLR 2024 · 被引用 61 次
