Free Energy Mixer
Jiecheng Lu, Shihao Yang
摘要
Standard attention stores keys/values losslessly but reads them via a per-head convex average, blocking channel-wise selection. We propose the Free Energy Mixer (FEM): a free-energy (log-sum-exp) read that applies a value-driven, per-channel log-linear tilt to a fast prior (e.g., from queries/keys in standard attention) over indices. Unlike methods that attempt to improve and enrich the scoring distribution, FEM treats it as a prior and yields a value-aware posterior read at unchanged complexity, smoothly moving from averaging to per-channel selection as the learnable inverse temperature increases, while still preserving parallelism and the original asymptotic complexity ( for softmax; for linearizable variants). We instantiate a two-level gated FEM that is plug-and-play with standard and linear attention, linear RNNs and SSMs. It consistently outperforms strong baselines on NLP, vision, and time-series at matched parameter budgets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
相关 Paper
- HyperMLP: An Integrated Perspective for Sequence ModelingJiecheng Lu, Shihao YangICML 2026 · 被引用 2 次
- Native Hybrid Attention for Efficient Sequence ModelingJusen Du, Jiaxi Hu, Zhang Tao, Weigao Sun 等ACL 2026 · 被引用 8 次
- Paramixer: Parameterizing Mixing Links in Sparse Factors Works Better than Dot-Product Self-AttentionTong Yu, Ruslan Khalitov, Lei Cheng, Zhirong YangCVPR 2022 · 被引用 11 次
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear TransformersKazuki Irie, Morris Yau, Samuel J. GershmanNeurIPS 2025 · 被引用 13 次
- Do We Really Need Complicated Model Architectures For Temporal Networks?Weilin Cong, Si Zhang, Jian Kang, Baichuan Yuan 等ICLR 2023 · 被引用 19 次
