MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
Zonglin Yang, Lidong Bing
摘要
While large language models (LLMs) show promise in scientific discovery, existing research focuses on inference or feedback-driven training, leaving the direct modeling of the generative reasoning process, (), unexplored. We demonstrate that directly training is mathematically intractable due to the combinatorial complexity () inherent in retrieving and composing inspirations from a vast knowledge base. To break this barrier, we introduce MOOSE-Star, a unified framework that enables tractable and scalable training of , while supporting more scalable inference. In the best case, MOOSE-Star reduces complexity from exponential to logarithmic () by (1) training on decomposed subtasks derived from the probabilistic equation of discovery, (2) employing motivation-guided hierarchical search to enable logarithmic retrieval and prune irrelevant subspaces, and (3) utilizing bounded composition for robustness against retrieval noise. To facilitate this, we release TOMATO-Star, a dataset of 108,717 decomposed papers (38,400 GPU hours) for training. Empirically, MOOSE-Star scales continuously with training data and inference budget, whereas direct brute-force sampling hits a "complexity wall."
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Reverse-Engineered Reasoning for Open-Ended GenerationHaozhe Wang, Haoran Que, Qixin Xu, Minghao Liu 等ICLR 2026 · 被引用 38 次
- SciMON: Scientific Inspiration Machines Optimized for NoveltyQingyun Wang, Doug Downey, Heng Ji, Tom HopeACL 2024 · 被引用 22 次
- MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical SearchZonglin Yang, Wanhao Liu, Ben Gao, Yujie Liu 等NeurIPS 2025 · 被引用 19 次
相关 Paper
- MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific HypothesesZonglin Yang, Wanhao Liu, Ben Gao, Tong Xie 等ICLR 2025 · 被引用 2 次
- Towards Multimodal Data-Driven Scientific Discovery Powered by LLM AgentsFan Liu, Xiaozhao Zeng, Hao LiuICLR 2026
- On the Emergence and Test-Time Use of Structural Information in Large Language ModelsMichelle Chao Chen, Moritz Miller, Bernhard Schölkopf, Siyuan GuoACL 2026 · 被引用 1 次
- HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget ReallocationFeng Xiong, Hongling Xu, Yifei Wang, Runxi Cheng 等EMNLP 2025 · 被引用 18 次
- Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement LearningGuanting Dong, Yifei Chen, Xiaoxi Li, Jiajie Jin 等SIGIR 2026 · 被引用 1 次
