Automated Hypothesis Validation with Agentic Sequential Falsifications
Kexin Huang, Ying Jin, Ryan Li, Michael Y. Li, Emmanuel J. Candès, Jure Leskovec
摘要
Hypotheses are central to information acquisition, decision-making, and discovery. However, many real-world hypotheses are abstract, highlevel statements that are difficult to validate directly. This challenge is further intensified by the rise of hypothesis generation from Large Language Models (LLMs), which are prone to hallucination and produce hypotheses in volumes that make manual validation impractical. Here we propose POPPER, an agentic framework for rigorous automated validation of free-form hypotheses. Guided by Karl Popper's principle of falsification, POPPER validates a hypothesis using LLM agents that design and execute falsification experiments targeting its measurable implications. A novel sequential testing framework ensures strict Type-I error control while actively gathering evidence from diverse observations, whether drawn from existing data or newly conducted procedures. We demonstrate POPPER on six domains including biology, economics, and sociology. POPPER delivers robust error control, high power, and scalability. Furthermore, compared to human scientists, POPPER achieved comparable performance in validating complex biological hypotheses while reducing time by 10 folds, providing a scalable, rigorous solution for hypothesis validation. POP-PER is freely available at https://github. com/snap-stanford/POPPER .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AutoDiscovery: Open-ended Scientific Discovery via Bayesian SurpriseDhruv Agarwal, Bodhisattwa Prasad Majumder, Reece Adamson, Megha Chakravorty 等NeurIPS 2025 · 被引用 35 次
- HeurekaBench: A Benchmarking Framework for AI Co-scientistSiba Smarak Panigrahi, Jovana Videnovic, Maria BrbicICLR 2026 · 被引用 10 次
- CausalGame: Benchmarking Causal Thinking of LLM Agents in GamesZhenhao Chen, Yongqiang Chen, Chenxi Liu, Junchi Yu 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper8
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Hypothesis Search: Inductive Reasoning with Language ModelsRuocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu 等ICLR 2024 · 被引用 156 次
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementLinlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar 等ICLR 2024 · 被引用 114 次
- Instruction Induction: From Few Examples to Natural Language Task DescriptionsOr Honovich, Uri Shaham, Samuel R. Bowman, Omer LevyACL 2023 · 被引用 48 次
- A large-scale benchmark for few-shot program induction and synthesisFerran Alet, Javier Lopez-Contreras, James Koppel, Maxwell I. Nye 等ICML 2021 · 被引用 20 次
相关 Paper
- Principle-Evolvable Scientific Discovery via Uncertainty MinimizationYingming Pu, Tao LIN, Hongyu ChenICML 2026 · 被引用 4 次
- SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI EnvironmentsSyed Yusuf Ahmed, Shiwei Feng, Chanwoo Bae, Calix Barrus 等ICSE 2026
- MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific HypothesesZonglin Yang, Wanhao Liu, Ben Gao, Tong Xie 等ICLR 2025 · 被引用 2 次
- Agentic Proposing: Enhancing Large language Model Reasoning via Compositional Skill SynthesisZhengbo Jiao, Shaobo Wang, Zifan Zhang, Xuan Ren 等ICML 2026 · 被引用 10 次
- STARK: Strategic Team of Agents for Refining KernelsJuncheng Dong, Yang Yang, Tao Liu, Yang Wang 等ICLR 2026 · 被引用 26 次
