LexTempus: Enhancing Temporal Generalizability of Legal Language Models Through Dynamic Mixture of Experts
T. Y. S. S. Santosh, Tuan-Quang Vuong
Abstract
The rapid evolution of legal concepts over time necessitates that legal language models adapt swiftly accounting for the temporal dynamics. However, prior works have largely neglected this crucial dimension, treating legal adaptation as a static problem rather than a continuous process. To address this gap, we pioneer Lex-Tempus, a dynamic mixture of experts model that explicitly models the temporal evolution of legal language in a parameter-efficient on-line learning framework. LexTempus starts with a single lightweight adapter expert and dynamically expands by adding new experts as significant deviations in the data distribution are detected. This self-expansion strategy allows LexTempus to adapt to new information without forgetting past knowledge, thereby improving temporal generalization. We use a a non-parametric similarity-based router to merge relevant experts into a unified expert for each test instance, ensuring efficient inference without additional overhead. We validate the effectiveness of LexTempus on ECHR and EU case law datasets, demonstrating its superiority in both perplexity and open-ended text generation quality metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9f16738-13b4-4309-b6b5-49ba5c1797a2Builds on32
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang et al.ICML 2024 · 605 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
Related papers
- Test-Time Mixture of World Models for Embodied Agents in Dynamic EnvironmentsJinwoo Jang, Minjong Yoo, Sihyung Yoon, Honguk WooICLR 2026 · 2 citations
- ChronosLex: Time-aware Incremental Training for Temporal Generalization of Legal Classification TasksT. Y. S. S. Santosh, Tuan-Quang Vuong, Matthias GrabmairACL 2024
- Dynamic TMoE: A Drift-Aware Dynamic Mixture of Experts Framework for Non-Stationary Time Series ForecastingJiawen Zhu, Shuhan Liu, Di Weng, Yingcai WuICML 2026
- Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-ExpertsXue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang et al.ACL 2025 · 9 citations
- Rewiring Experts on the Fly: Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert ModelsGuinan Su, Yanwu Yang, Li Shen, Lu Yin et al.ICML 2026 · 3 citations
