L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
Samuel Holt, Max Ruiz Luyten, Mihaela van der Schaar
摘要
Transformer-based large language models (LLMs) are constrained by the fixed context window of the underlying transformer architecture, hindering their ability to produce long and coherent outputs. Memory-augmented LLMs are a promising solution, but current approaches cannot handle long output generation tasks since they (1) only focus on reading memory and reduce its evolution to the concatenation of new memories or (2) use very specialized memories that cannot adapt to other domains. This paper presents L2MAC 1 , the first practical LLM-based general-purpose stored-program automatic computer (von Neumann architecture) framework, an LLM-based multi-agent system, for long and consistent output generation. Its memory has two components: the instruction registry, which is populated with a prompt program to solve the user-given task, and a file store, which will contain the final and intermediate outputs. Each instruction in turn is executed by a separate LLM agent, whose context is managed by a control unit capable of precise memory reading and writing to ensure effective interaction with the entire file store. These components enable L2MAC to generate extensive outputs, bypassing the constraints of the finite context window while producing outputs that fulfill a complex user-specified task. We empirically demonstrate that L2MAC achieves state-of-the-art performance in generating large codebases for system design tasks, significantly outperforming other coding methods in implementing the detailed user-specified task; we show that L2MAC works for general-purpose extensive text-based tasks, such as writing an entire book; and we provide valuable insights into L2MAC's performance improvement over existing methods. Published as a conference paper at ICLR 2024 app.py from flask import Flask, request, render_template from user import User from mock_db import MockDB from contact import Contact from message import Message from group_chat import GroupChat app = Flask(name) db = MockDB() @app.route('/') def home(): return render_template('home.html'), 200 @app.route('/register', methods=['POST']) def register(): data = request.get_json() user = User(data['email'], data['password']) contact = Contact(data['email']) db.add(user.email, user) db.add(contact.email, contact) return render_template('register.html', message='User registered successfully'), 200 @app.route('/login', methods=['POST']) def login(): data = request.get_json() user = db.retrieve(data['email']) if user and user.password == data['password']: user.set_online_status(True) return render_template('login.html', message='Login successful'), 200 else: return render_template('login.html', message='Invalid email or password'), 401 @app.route('/logout', methods=['POST']) def logout(): data = request.get_json() user = db.retrieve(data['email']) if user and user.password == data['password']: user.set_online_status(False) return 'Logout successful', 200 else: return 'Invalid email or password', 401 @app.route('/send_message', methods=['POST']) def send_message(): data = request.get_json() sender = db.retrieve(data['sender']) receiver = db.retrieve(data['receiver']) if sender and receiver and sender.password == data['password']: message = Message(sender.email, receiver.email, data['content']) if receiver.get_online_status(): receiver.receive_message(message.send_message()) else: receiver.add_to_message_queue(message.send_message()) return 'Message sent', 200 else: return 'Invalid email or password', 401 if name == 'main': app.run(debug=True) contact.py class Contact: def init(self, email): self.email = email self.blocked_contacts = set() self.groups = def block_contact(self, contact_email): self.blocked_contacts.add(contact_email) 1. Setup structure 2. Create registration Contributions: 1 ⃝ We introduce the L2MAC framework, the first practical LLM-based generalpurpose stored-program automatic computer (von Neumann architecture) framework, an LLM-based multi-agent system, for long output generation tasks. 2 ⃝ We provide a practical implementation of this framework for code generation tasks called Code-L2MAC. This uses a Control Unit to control the input and output of the LLM and the use of entire memory file store read/write tools and highlights Published as a conference paper at ICLR 2024 Invoke Tool? Called Step Completed? 1. Summarize step output to Mrs and clear Ct. 2. Load next I(k)aa, and start new Ct=I(k),Mrs Append cycle message Mc Would exceed context? 1. Unwind oldest messages M, till 2. Summarize progress to Mrs and clear Ct. 3. Restart same I(k), with Ct=I(k),Mrs
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Discovering Preference Optimization Algorithms with and for Large Language ModelsChris Lu, Samuel Holt, Claudio Fanconi, Alex J. Chan 等NeurIPS 2024 · 被引用 41 次
- Automatically Learning Hybrid Digital Twins of Dynamical SystemsSamuel Holt, Tennison Liu, Mihaela van der SchaarNeurIPS 2024 · 被引用 26 次
- Data-Driven Discovery of Dynamical Systems in Pharmacology using Large Language ModelsSamuel Holt, Zhaozhi Qian, Tennison Liu, James Weatherall 等NeurIPS 2024 · 被引用 16 次
- COFFE: A Code Efficiency Benchmark for Code GenerationYun Peng, Jun Wan, Yichen Li, Xiaoxue RenFSE 2025 · 被引用 8 次
- QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language ModelsQirui Zhou, Yuanbo Wen, Ruizhi Chen, Ke Gao 等AAAI 2025 · 被引用 7 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- OMAC: A Holistic Optimization Framework for LLM-Based Multi-Agent CollaborationShijun Li, Hilaf Hasson, Joydeep GhoshICML 2026
- Memory OS of AI AgentJiazheng Kang, Mingming Ji, Zhe Zhao, Ting BaiEMNLP 2025 · 被引用 4 次
- Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular DataFengxian Dong, Zhi Zheng, Xiao Han, Wei Chen 等ACL 2026
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsYi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan 等ACL 2026 · 被引用 40 次
- RepLLM: Toward Automatically Reproducing Network Research ResultsYining Jiang, Yunxin Xu, Wenyun Xu, Yufan Zhu 等SIGCOMM 2026
