Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
Yulu Gan, Phillip Isola
Abstract
Large model Small model Gaussian search window tasks GSM8k expert (math) ROCStories expert (writing) MBPP expert (Programming) USPTO expert (chemistry) (a) (b) (c) Post-training with RandOpt O(1) training and FLOP-efficient with Better Acc Scaling Law Figure 1: (a) Schematic of the main effects we observe (see Fig 2 for a version with real data). Left: Small models live in a needle in a haystack regime, where good solutions to downstream tasks occupy a tiny fraction of the surrounding weights. In this regime, it is important to have a smart search algorithm, such as gradient descent or other forms of iterative optimization. Right: Large models are surrounded by a veritable thicket of task-specific solutions. In this regime, random sampling is sufficient to quickly land on promising adaptations, which can then be ensembled to yield strong behavior, an approach we call RandOpt. (b) Solution density -i.e. density of task-improving weights in a Gaussian neighborhood of the pretrained weights -scales with model size. (c) RandOpt is O(1) in training steps, FLOP-efficient, and competitive in converged accuracy with GRPO and ES. Results are shown on the Countdown task with Olmo-3-7B-Instruct; RandOpt uses 5000 random weight guesses and ensembles the top K; K-pass baselines use Test-time Majority Vote (TT-MV). More results are shown in Fig. 6 and Table 4 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f82ffd6-492d-4ab2-9dbc-a1f7ba325e6aCited by top-tier papers1
Ask how each one uses itBuilds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
Related papers
- Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization AlignmentChenghao Fan, Zhenyi Lu, Sichen Liu, Chengfeng Gu et al.ICML 2025
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for ReasoningXuechen Zhang, Zijian Huang, Yingcong Li, Chenshun Ni et al.NeurIPS 2025 · 66 citations
- EGSS: Entropy-guided Stepwise Scaling for Reliable Software EngineeringChenhui Mao, Yuanting Lei, Zhixiang Wei, Ming Liang et al.ACL 2026
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive ExplorationZhicheng Yang, Zhijiang Guo, Yinya Huang, Yongxin Wang et al.ICML 2026 · 38 citations
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsZhiyuan Liang, Dongwen Tang, Yuhao Zhou, Xuanlei Zhao et al.NeurIPS 2025 · 22 citations
