Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
Yulu Gan, Phillip Isola
摘要
Large model Small model Gaussian search window tasks GSM8k expert (math) ROCStories expert (writing) MBPP expert (Programming) USPTO expert (chemistry) (a) (b) (c) Post-training with RandOpt O(1) training and FLOP-efficient with Better Acc Scaling Law Figure 1: (a) Schematic of the main effects we observe (see Fig 2 for a version with real data). Left: Small models live in a needle in a haystack regime, where good solutions to downstream tasks occupy a tiny fraction of the surrounding weights. In this regime, it is important to have a smart search algorithm, such as gradient descent or other forms of iterative optimization. Right: Large models are surrounded by a veritable thicket of task-specific solutions. In this regime, random sampling is sufficient to quickly land on promising adaptations, which can then be ensembled to yield strong behavior, an approach we call RandOpt. (b) Solution density -i.e. density of task-improving weights in a Gaussian neighborhood of the pretrained weights -scales with model size. (c) RandOpt is O(1) in training steps, FLOP-efficient, and competitive in converged accuracy with GRPO and ES. Results are shown on the Countdown task with Olmo-3-7B-Instruct; RandOpt uses 5000 random weight guesses and ensembles the top K; K-pass baselines use Test-time Majority Vote (TT-MV). More results are shown in Fig. 6 and Table 4 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
相关 Paper
- Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization AlignmentChenghao Fan, Zhenyi Lu, Sichen Liu, Chengfeng Gu 等ICML 2025
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for ReasoningXuechen Zhang, Zijian Huang, Yingcong Li, Chenshun Ni 等NeurIPS 2025 · 被引用 66 次
- EGSS: Entropy-guided Stepwise Scaling for Reliable Software EngineeringChenhui Mao, Yuanting Lei, Zhixiang Wei, Ming Liang 等ACL 2026
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive ExplorationZhicheng Yang, Zhijiang Guo, Yinya Huang, Yongxin Wang 等ICML 2026 · 被引用 38 次
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsZhiyuan Liang, Dongwen Tang, Yuhao Zhou, Xuanlei Zhao 等NeurIPS 2025 · 被引用 22 次
