Regress, Don't Guess: A Regression-like Loss on Number Tokens for Language Models
Jonas Zausinger, Lars Pennig, Anamarija Kozina, Sean Sdahl, Julian Sikora, Adrian Dendorfer, Timofey Kuznetsov, Mohamad Hagog, Nina Wiedemann, Kacper Chlodny, Vincent Limbach, Anna Ketteler
Abstract
While language models have exceptional capabilities at text generation, they lack a natural inductive bias for emitting numbers and thus struggle in tasks involving quantitative reasoning, especially arithmetic. One fundamental limitation is the nature of the cross-entropy (CE) loss, which assumes a nominal scale and thus cannot convey proximity between generated number tokens. In response, we here present a regression-like loss that operates purely on token level. Our proposed Number Token Loss (NTL) comes in two flavors and minimizes either the L p norm or the Wasserstein distance between the numerical values of the real and predicted number tokens. NTL can easily be added to any language model and extend the CE objective during training without runtime overhead. We evaluate the proposed scheme on various mathematical datasets and find that it consistently improves performance in math-related tasks. In a direct comparison on a regression task, we find that NTL can match the performance of a regression head, despite operating on token level. Finally, we scale NTL up to 3B parameter models and observe improved performance, demonstrating its potential for seamless integration into LLMs. We hope to inspire LLM developers to improve their pretraining objectives and distribute NTL as a minimalistic and lightweight PyPI package ntloss: https://github.com/ai4sd/number- token-loss. Development code for full paper reproduction is available separately.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- FoNE: Precise Single-Token Number Embeddings via Fourier FeaturesTianyi Zhou, Deqing Fu, Mahdi Soltanolkotabi, Robin Jia et al.ICLR 2026 · 24 citations
- LLM2Fx-Tools: Tool Calling for Music Post-ProductionSeungHeon Doh, Junghyun Koo, Marco A. Martínez-Ramírez, Woosung Choi et al.ICLR 2026 · 10 citations
- Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement LearningMing Chen, Sheng Tang, Rong-Xi Tan, Ziniu Li et al.ICML 2026 · 2 citations
- CONE: Embeddings for Complex Numerical Data Preserving Unit and Variable SemanticsGyanendra Shrestha, Anna Pyayt, Michael N. GubanovSIGMOD 2026
- Enhancing Numerical Prediction in LLMs via Smooth MMD AlignmentZhuo Zuo, Li Yue, Wenhao Zheng, Chenpeng Wang et al.ICML 2026
Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 163 citations
Related papers
- Enhancing Numerical Prediction of MLLMS With Soft LabelingPei Wang, Zhaowei Cai, Hao Yang, Davide Modolo et al.ICCV 2025 · 1 citation
- GeoNum: Bridging Numerical Continuity and Language Semantics via Geometric EmbeddingShengkai Jin, Tianyu Chen, Chonghan Gao, Jun HanAAAI 2026
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman et al.ICLR 2026 · 8 citations
- Efficient numeracy in language models through single-token number embeddingsLinus Kreitner, Paul Hager, Jonathan Mengedoht, Georgios Kaissis et al.ICML 2026 · 5 citations
- Methods for Numeracy-Preserving Word EmbeddingsDhanasekar Sundararaman, Shijing Si, Vivek Subramanian, Guoyin Wang et al.EMNLP 2020 · 28 citations
