Lune

ICML2026Top-tier venue

Cycle-of-Science: Reliable Reasoning through Counterfactual Verification for Agent Decision Making

Ruojie Zhang, Wencheng Zhu, Peiyuan Jiang, dayong zhu

2026Year

Abstract

Large Language Models have significantly advanced autonomous agents through their sophisticated perception and execution capabilities. Despite effective, agents still struggle with robust decision-making due to passive learning from similar experiences that often confound correlation with causality. Inspired by the Scientific Method, we propose a Cycle-of-Science framework that autonomously explores potential causal pathways through an iterative loop of Hypothesis, Experiment, and Validation, enabling agents to identify truly effective causal dependencies. To be specific, we first leverage causal knowledge to guide the initial hypotheses generation. These hypotheses are then analyzed through experiments using counterfactual samples. Afterward, we perform causal analysis to quantify effects of interventions, deriving well-validated hypotheses for next agent steps. To train our policy, we further introduce a two-stage pipeline that integrates supervised fine-tuning with Counterfactual Preference Optimization, which constructs preference signals from intervention outcomes to reinforce validated reasoning chains. Experiments on benchmarks demonstrate that our method achieves superior performance over state-of-the-art approaches.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 418b7571-a3ba-41d5-b30f-67f0adc3bc97

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines