Fourier Head: Helping Large Language Models Learn Complex Probability Distributions
Nate Gillman, Daksh Aggarwal, Michael Freeman, Chen Sun
Abstract
As the quality of large language models has improved, there has been increased interest in using them to model non-linguistic tokens. For example, the Decision Transformer recasts agentic decision making as a sequence modeling problem, using a decoder-only LLM to model the distribution over the discrete action space for an Atari agent. However, when adapting LLMs to non-linguistic domains, it remains unclear if softmax over discrete bins captures the continuous structure of the tokens and the potentially complex distributions needed for high quality token generation. We introduce a neural network layer, constructed using Fourier series, which we can easily substitute for any linear layer if we want the outputs to have a more continuous structure. We perform extensive analysis on synthetic datasets, as well as on large-scale decision making and time series forecasting tasks. We also provide theoretical evidence that this layer can better learn signal from data while ignoring high-frequency noise. All of our results support the effectiveness of our proposed Fourier head in scenarios where the underlying data distribution has a natural continuous structure. For example, the Fourier head improves a Decision Transformer agent's returns by 46% on the Atari Seaquest game, and increases a state-of-the-art times series foundation model's forecasting performance by 3.5% across 20 benchmarks unseen during training. We release our implementation at https://nategillman.com/fourier-head . Fourier Head Learns Higher Quality Densities Figure 1 : We task an MLP with learning to approximate a continuous bimodal density using a categorical distribution and a cross entropy objective. We observe that a standard linear classification head fails to distinguish between the two modes, and overfits to high-frequency noise in the training set. In contrast, our proposed Fourier head learns a smoother, more accurate categorical distribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a1d9c58-b34a-4e8e-be6a-5beb912f339aCited by top-tier papers4
- Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement LearningWenchang Duan, Yaoliang Yu, Jiwan He, Yi ShiNeurIPS 2025 · 11 citations
- D-Models and E-Models: Diversity-Stability Trade-offs in the Sampling Behavior of Large Language ModelsJia Gu, Liang Pang, Huawei Shen, Xueqi ChengWWW 2026
- Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length GeneralizationErmo Hua, Che Jiang, Xingtai Lv, Kaiyan Zhang et al.ICML 2025
- Text-to-Distribution Prediction with Quantile Tokens and Neighbor ContextYilun Zhu, Yuan Zhuang, Nikhita Vedula, Dushyanta Dhyani et al.ACL 2026
Builds on6
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 601 citations
Related papers
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
- Language Models Are Implicitly ContinuousSamuele Marro, Davide Evangelista, Xuanqiang Angelo Huang, Emanuele La Malfa et al.ICLR 2025
- From Tokenizer Bias to Backbone Capability: A Controlled Study of LLMs for Time Series ForecastingXinyu Zhang, Shanshan Feng, Xutao Li, Kenghong Lin et al.KDD 2026 · 2 citations
- Rethinking Fourier Transform from A Basis Functions Perspective for Long-term Time Series ForecastingRunze Yang, Longbing Cao, Jie Yang, Jianxun LiNeurIPS 2024 · 37 citations
- Reasoning is Periodicity? Improving Large Language Models Through Effective Periodicity ModelingYihong Dong, Ge Li, Xue Jiang, Yongding Tao et al.NeurIPS 2025 · 1 citation
