HistLens: Mapping Idea Change across Concepts and Corpora
Yi Jing, Weiyun Qiu, Yihang Peng, Zhifang Sui
Abstract
Language change both reflects and shapes social processes, and the semantic evolution of foundational concepts provides a measurable trace of historical and social transformation. Despite recent advances in diachronic semantics and discourse analysis, existing computational approaches often (i) concentrate on a single concept or a single corpus, making findings difficult to compare across heterogeneous sources, and (ii) remain confined to surface lexical evidence, offering insufficient computational and interpretive granularity when concepts are expressed implicitly. We propose HistLens, a unified, SAE-based framework for multi-concept, multi-corpus conceptual-history analysis. The framework decomposes concept representations into interpretable features and tracks their activation dynamics over time and across sources, yielding comparable conceptual trajectories within a shared coordinate system. Experiments on long-span press corpora show that HistLens supports cross-concept, cross-corpus computation of patterns of idea evolution and enables implicit concept computation. By bridging conceptual modeling with interpretive needs, HistLens broadens the analytical perspectives and methodological repertoire available to social science and the humanities for diachronic text analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5cbcc2e-9a32-4935-9820-8ca10c969c6eBuilds on6
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational QualityWilliam Timkey, Marten van SchijndelEMNLP 2021 · 59 citations
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh et al.ICLR 2025 · 10 citations
- LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-EncoderYi Jing, Zijun Yao, Hongzhu Guo, Lingxu Ran et al.EMNLP 2025 · 7 citations
- Media Framing: A typology and Survey of Computational Approaches Across DisciplinesYulia Otmakhova, Shima Khanehzar, Lea FrermannACL 2024 · 2 citations
Related papers
- Word2Fun: Modelling Words as Functions for Diachronic Word RepresentationBenyou Wang, Emanuele Di Buccio, Massimo MelucciNeurIPS 2021 · 6 citations
- Analysing Lexical Semantic Change with Contextualised Word RepresentationsMario Giulianelli, Marco Del Tredici, Raquel FernándezACL 2020 · 118 citations
- Modeling the Evolution of English Noun Compounds with Feature-Rich Diachronic Compositionality PredictionFilip Miletic, Sabine Schulte im WaldeACL 2025
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsBenno Krojer, Perampalli Shravan Nayak, Oscar Mañas, Vaibhav Adlakha et al.ICML 2026 · 6 citations
- Interpretable Word Sense Representations via Definition Generation: The Case of Semantic Change AnalysisMario Giulianelli, Iris Luden, Raquel Fernández, Andrey KutuzovACL 2023 · 9 citations
