Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
DiJia Su, Sainbayar Sukhbaatar, Michael Rabbat, Yuandong Tian, Qinqing Zheng
摘要
In cognition theory, human thinking is governed by two systems: the fast and intuitive System 1 and the slower but more deliberative System 2. Analogously, Large Language Models (LLMs) can operate in two reasoning modes: outputting only the solutions (fast mode) or both the reasoning chain and the final solution (slow mode). We present Dualformer, a single Transformer model that seamlessly integrates both the fast and slow reasoning modes by training on randomized reasoning traces, where different parts of the traces are strategically dropped during training. At inference time, Dualformer can be easily configured to execute in either fast or slow mode, or automatically decide which mode to engage (auto mode). It outperforms baselines in both performance and computational efficiency across all three modes: (1) in slow mode, Dualformer achieves 97.6% optimal rate on unseen 30 × 30 maze tasks, surpassing the Searchformer baseline (93.3%) trained on data with complete reasoning traces, with 45.5% fewer reasoning steps; (2) in fast mode, Dualformer achieves 80% optimal rate, significantly outperforming the Solution-Only model trained on solution-only data, which has an optimal rate of only 30%; (3) in auto mode, Dualformer achieves 96.6% optimal rate with 59.9% fewer steps than Searchformer. Moreover, Dualformer produces more diverse reasoning traces than Searchformer. For math reasoning problems, our techniques have also achieved improved performance with LLM fine-tuning, demonstrating its generalization beyond task-specific models. We open source our code at https://github.com/facebookresearch/dualformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement LearningWenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang 等NeurIPS 2025 · 被引用 85 次
- ARM: Adaptive Reasoning ModelSiye Wu, Jian Xie, Yikai Zhang, Aili Chen 等NeurIPS 2025 · 被引用 31 次
- Controlling Thinking Speed in Reasoning ModelsZhengkai Lin, Zhihang Fu, Ze Chen, Chao Chen 等NeurIPS 2025 · 被引用 20 次
- LIMOPro: Reasoning Refinement for Efficient and Effective Test-time ScalingYang Xiao, Jiashuo Wang, Ruifeng Yuan, Chunpu Xu 等NeurIPS 2025 · 被引用 15 次
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for ReasoningAntonios Valkanas, Soumyasundar Pal, Pavel Rumiantsev, Yingxue Zhang 等NeurIPS 2025 · 被引用 10 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
相关 Paper
- Thinker: Learning to Think Fast and SlowStephen Chung, Wenyu Du, Jie FuNeurIPS 2025 · 被引用 10 次
- Decoupling Knowledge and Reasoning in LLMs: An Exploration Using Cognitive Dual-System TheoryMutian Yang, Jiandong Gao, Ji WuAAAI 2026 · 被引用 5 次
- PRIME: Planning and Retrieval-Integrated Memory for Enhanced ReasoningHieu Tran, Zonghai Yao, Nguyen Luong Tran, Zhichao Yang 等AAAI 2026 · 被引用 1 次
- Teach Small Models to Reason by Curriculum DistillationWangyi Jiang, Yaojie Lu, Hongyu Lin, Xianpei Han 等EMNLP 2025
- Think Only When You Need with Large Hybrid-Reasoning ModelsLingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong 等NeurIPS 2025 · 被引用 71 次
