Lune

ICML2026顶会

MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier

Zonglin Yang, Lidong Bing

2026年份

摘要

While large language models (LLMs) show promise in scientific discovery, existing research focuses on inference or feedback-driven training, leaving the direct modeling of the generative reasoning process, P(hypothesis∣background)P(\text{hypothesis}|\text{background}) (P(h∣b)P(h|b)), unexplored. We demonstrate that directly training P(h∣b)P(h|b) is mathematically intractable due to the combinatorial complexity (O(Nk)O(N^k)) inherent in retrieving and composing inspirations from a vast knowledge base. To break this barrier, we introduce MOOSE-Star, a unified framework that enables tractable and scalable training of P(h∣b)P(h|b), while supporting more scalable inference. In the best case, MOOSE-Star reduces complexity from exponential to logarithmic (O(log⁡N)O(\log N)) by (1) training on decomposed subtasks derived from the probabilistic equation of discovery, (2) employing motivation-guided hierarchical search to enable logarithmic retrieval and prune irrelevant subspaces, and (3) utilizing bounded composition for robustness against retrieval noise. To facilitate this, we release TOMATO-Star, a dataset of 108,717 decomposed papers (38,400 GPU hours) for training. Empirically, MOOSE-Star scales continuously with training data and inference budget, whereas direct brute-force sampling hits a "complexity wall."

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 8187d9c7-499a-4e72-b69f-9d1150ddc0ff

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖