Free Energy Mixer
Jiecheng Lu, Shihao Yang
Abstract
Standard attention stores keys/values losslessly but reads them via a per-head convex average, blocking channel-wise selection. We propose the Free Energy Mixer (FEM): a free-energy (log-sum-exp) read that applies a value-driven, per-channel log-linear tilt to a fast prior (e.g., from queries/keys in standard attention) over indices. Unlike methods that attempt to improve and enrich the scoring distribution, FEM treats it as a prior and yields a value-aware posterior read at unchanged complexity, smoothly moving from averaging to per-channel selection as the learnable inverse temperature increases, while still preserving parallelism and the original asymptotic complexity ( for softmax; for linearizable variants). We instantiate a two-level gated FEM that is plug-and-play with standard and linear attention, linear RNNs and SSMs. It consistently outperforms strong baselines on NLP, vision, and time-series at matched parameter budgets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- HyperMLP: An Integrated Perspective for Sequence ModelingJiecheng Lu, Shihao YangICML 2026 · 2 citations
- Native Hybrid Attention for Efficient Sequence ModelingJusen Du, Jiaxi Hu, Zhang Tao, Weigao Sun et al.ACL 2026 · 8 citations
- Paramixer: Parameterizing Mixing Links in Sparse Factors Works Better than Dot-Product Self-AttentionTong Yu, Ruslan Khalitov, Lei Cheng, Zhirong YangCVPR 2022 · 11 citations
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear TransformersKazuki Irie, Morris Yau, Samuel J. GershmanNeurIPS 2025 · 13 citations
- Do We Really Need Complicated Model Architectures For Temporal Networks?Weilin Cong, Si Zhang, Jian Kang, Baichuan Yuan et al.ICLR 2023 · 19 citations
