Underlying Dimensions in Semantic Change

Lexical semantic change detection is a well-established field in computational linguistics and natural language processing, with several shared tasks and curated datasets (e.g., Basile et al, 2020; Kutuzov & Pivovarova, 2021; Schlechtweg et al, 2020). Modern approaches in the study of semantic change often consist of applications of vector space modelling (VSM) (Tahmasebi et al, 2021; Tahmasebi & Dubossarsky, 2023), which converts raw text into numerical vector representations then quantifies changes through clustering methods or distance measures (e.g., cosine distance) to those vectors.

 A key conceptual problem of the above shared tasks and models is that they operationalise ‘semantic change’ as change in the distribution of senses of words (between 2 time periods). Additionally, these approaches tend to suffer from the sensitivity to corpus size (Antoniak & Mimno, 2018; Sahlgren & Lenci, 2016) and the lack of interpretability (Lenci et al, 2022).

In practice, many types of change do not map neatly onto such clear-cut sense distinctions. Subtle shifts in typical collocates, prototypical structures (Geeraerts, 1997), argument-structure (Christiansen & Joseph, 2016), or conceptual frame (Traugott, 2017) may leave a coarse-grained sense inventory intact while still reflecting substantial reorganization in how a word is used and conceptualised. For instance, meeting retains the general ORGANISED GATHERING meaning, yet its usage profile develops a digitally mediated default, with new collocates like Zoom/Teams, virtual, hybrid. A strictly sense-based operationalisation therefore risks underestimating these more graded, profile-level developments, which motivates approaches that track changes in usage profiles directly rather than inferring them only through sense annotations.

We present an approach that 1) is conceptually simpler, 2) requires less data than VSM methods, 3) is more direct and theory-driven in interpreting the changes across different dimensions of words’ usages and meanings.

Given a concept/lexeme, we 1) annotate the different (theory-driven) dimensions of word usages using available tools such as automatic parsers or Large Language Models (LLMs), 2) measure change in these dimensions using divergence measures, 3) interpret the changing dimensions through divergence decomposition.

To demonstrate the feasibility of this approach, we report 3 case studies: 

In the 2 case studies with SemEval dataset (Phan-Tat et al, 2026b, 2026a), we demonstrate that it can be used for lexical semantic change detection and can outperform many VSM approaches in many cases while remaining lightweight and transparent. The models are demonstrated to capture semantically relevant signals and supports interpretable explanations of the observed changes. 

In Case Study 3, we applied the framework to three chemical concepts from late eighteenth- to early nineteenth-century chemistry and demonstrated that it successfully captures the linguistic imprint of a major paradigm shift, the emergence of oxygen chemistry. We worked with a small and highly imbalanced subset of the Royal Society Corpus 6.0 Open (Fischer et al, 2020), which is challenging for many VSM approaches. 

Because the method is lightweight and relies only on off-the-shelf AI/NLP tools, it can be applied to diverse diachronic corpora and domains. Concretely, it identifies which dimensions shift, when the shifts occur, and which lexical items contribute most to the observed divergence, thereby providing testable targets and constraints for modelling assumptions and simulations.

Acknowledgements

This work was supported by the Horizon Europe MSCA Doctoral Network project CASCADE (101119511) (www.horizoncascade.net).

References 

Antoniak, M., & Mimno, D. (2018). Evaluating the Stability of Embedding based Word Similarities. Transactions of the Association for Computational Linguistics, 6, 107–119.

Basile, P., Caputo, A., Caselli, T., Cassotti, P., & Varvara, R. (2020). DIACR-Ita @ EVALITA2020: Overview of the EVALITA2020 Diachronic Lexical Semantics (DIACR-Ita) Task. In V. Basile, D. Croce, M. Maro, & L. C. Passaro (Eds.), EVALITA Evaluation of NLP and Speech Tools for Italian - December 17th, 2020 (pp. 411–419). Torino: Accademia University Press.

Christiansen, B. J., & Joseph, B. D. (2016). On the relationship between argument structure change and semantic change. Proceedings of the Linguistic Society of America, 1, 26.

Fischer, S., Knappen, J., Menzel, K., & Teich, E. (2020). The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study.

Geeraerts, D. (1997). Diachronic Prototype Semantics: A Contribution Historical Lexicology. Oxford University Press.

Kutuzov, A., & Pivovarova, L. (2021). RuShiftEval: A Shared Task on Semantic Shift Detection for Russian. In Computational linguistics and intellectual technologies.

Lenci, A., Sahlgren, M., Jeuniaux, P., Cuba Gyllensten, A., & Miliani, M. (2022). A comparative evaluation and analysis of three generations of Distributional Semantic Models. Language Resources and Evaluation, 56(4), 1269–1313.

Phan-Tat, B., Heylen, K., Geeraerts, D., Pascale, S. D., & Speelman, D. (2026a). Reframe or remain: Unsupervised lexical semantic change detection with frame semantics.

Phan-Tat, B., Heylen, K., Geeraerts, D., Pascale, S. D., & Speelman, D. (2026b). Transparent semantic change detection with dependency-based profiles.

Sahlgren, M., & Lenci, A. (2016). The Effects of Data Size and Frequency Range on Distributional Semantic Models. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (pp. 975–980). Austin, Texas: Association for Computational Linguistics.

Schlechtweg, D., McGillivray, B., Hengchen, S., Dubossarsky, H., & Tahmasebi, N. (2020). SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection. In Proceedings of the Fourteenth Workshop on Semantic Evaluation (pp. 1–23). Barcelona (online): International Committee for Computational Linguistics.

Tahmasebi, N., Borin, L., Jatowt, A., Xu, Y., & Hengchen, S. (2021). Computational approaches to semantic change.

Tahmasebi, N., & Dubossarsky, H. (2023). Computational modeling of semantic change. ( eprint: 2304.06337)

Traugott, E. C. (2017, March). Semantic Change. Oxford University Press.