In-Context Universal Approximation, Compositional Generalization, and Algorithm Emulation
Jerry Yao-Chieh Hu, Hong-Yu Chen, Po-Chiao Lin, Maojiang Su, Han Liu
Abstract
We study in-context universal approximation and compositional generalization in frozen softmax Transformers as prompt-programmable computation. We prove in-context universality via in-context emulation: a fixed-weight Transformer emulates target computations specified by the prompt and hence approximates a broad class of continuous sequence-to-sequence functions. Building on this view, we establish one-pass and multi-pass composition theorems: prompts associated with simple ``subprograms'' let the same fixed Transformer execute their composition and thereby synthesize more complex programs on-the-fly. These results support a principled view of prompts as programs and fixed-weight Transformers as program interpreters. They also provide concrete mechanisms by which GPT-style models execute and assemble algorithms in context. Please see arXiv for the full version.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on39
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
Related papers
- In-Context Algorithm Emulation in Fixed-Weight TransformersJerry Yao-Chieh Hu, Hude Liu, Jennifer Yuntong Zhang, Han LiuICLR 2026 · 7 citations
- Universal Approximation with Softmax AttentionJerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen, Weimin Wu et al.ICML 2026 · 9 citations
- Looped Transformers as Programmable ComputersAngeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee et al.ICML 2023 · 175 citations
- Prompting a Pretrained Transformer Can Be a Universal ApproximatorAleksandar Petrov, Philip Torr, Adel BibiICML 2024 · 19 citations
- Universal Length Generalization with Turing ProgramsKaiying Hou, David Brandfonbrener, Sham M. Kakade, Samy Jelassi et al.ICML 2025
