Doing Experiments and Revising Rules with Natural Language and Probabilistic Reasoning
Top Piriyakulkij, Cassidy Langenfeld, Tuan Anh Le, Kevin Ellis
Abstract
We give a model of how to infer natural language rules by doing experiments. The model integrates Large Language Models (LLMs) with Monte Carlo algorithms for probabilistic inference, interleaving online belief updates with experiment design under information-theoretic criteria. We conduct a human-model comparison on a Zendo-style task, finding that a critical ingredient for modeling the human data is to assume that humans also consider fuzzy, probabilistic rules, in addition to assuming that humans perform approximately-Bayesian belief updates. We also compare with recent algorithms for using LLMs to generate and revise hypotheses, finding that our online inference method yields higher accuracy at recovering the true underlying rule, and provides better support for designing optimal experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4d259d6-7286-4630-b2d2-ef9a29d21a07Cited by top-tier papers7
- Is Programming by Example Solved by LLMs?Wen-Ding Li, Kevin EllisNeurIPS 2024 · 45 citations
- PoE-World: Compositional World Modeling with Products of Programmatic ExpertsTop Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller et al.NeurIPS 2025 · 31 citations
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang et al.NeurIPS 2025 · 30 citations
- Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like PeopleGabriel Grand, Valerio Pepe, Joshua B. Tenenbaum, Jacob AndreasICLR 2026 · 7 citations
- Program Synthesis via Test-Time TransductionKang-il Lee, Jahyun Koo, Seunghyun Yoon, Minbeom Kim et al.NeurIPS 2025 · 4 citations
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 651 citations
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
Related papers
- Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on InferencesThomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger P. LevyEMNLP 2025
- Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic InferenceJiazheng Li, Hanqi Yan, Yulan HeACL 2025
- Bayesian Social Deduction with Graph-Informed Language ModelsShahab Rahimirad, Guven Gergerli, Lucia Romero, Angela Qian et al.ACL 2026 · 4 citations
- Language and Experience: A Computational Model of Social Learning in Complex TasksCédric Colas, Tracey Mills, Ben Prystawski, Michael Henry Tessler et al.ICLR 2026 · 1 citation
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell et al.EMNLP 2023 · 57 citations
