Session 4

Wednesday 15:00 - 16:30

High Tor 2

Chair: James Chetwood

Tracing progress: A diachronic lexical semantic approach to conceptual change in eighteenth-century Britain

KU Leuven

 Diachronic lexical semanticists and intellectual historians alike have long theorised about the mechanisms of conceptual change (Geeraerts, 1997; Koselleck, 2004; Kuukkanen, 2008). Specifically, in diachronic lexical semantics, prototype theory models semantic change by positing that lexical categories have central, highly representative cases, and more peripheral ones (Geeraerts et al., 2024) that shift and change places over time. Similarly, in the Begriffsgeschichte tradition, concepts are seen as having an invariable ‘core’ and a variable ‘margin’ (Kuukkanen, 2008). Despite these conceptual similarities, reflections have been conducted largely in parallel. This paper brings them into dialogue through a quantitative case study of the concept of progress in eighteenth-century Britain.

 The concept of progress has been of extensive interest for intellectual historians and philosophers (e.g., Nisbet, 2017; Spadafora, 1990; Wagner, 2016). Our analysis focuses on how progress changed from being a field-specific notion (e.g., progress in dramatic writing) to having a civilisation-wide conception (e.g., the progress of man) within the eighteenth century, as suggested by Spadafora (1990). We analyse syntactic constructions and collocational patterns that mark the transition between senses, reflecting the intellectual undertones of eighteenth-century Britain. We also ask (i) whether specific salient voices influenced this change and (ii) whether any identifiable areas of knowledge served as loci or domains of change.

Methodologically, we approach the problem from both the onomasiological and semasiological perspectives. To do this, we combine recent advances in distributional semantics (Geeraerts et al., 2024) with well-established corpus-based quantitative modelling techniques. To ensure relevance to intellectual history as well as machine readability, we use the Eighteenth Century Collections Online, transcribed by the Text Creation Partnership (ecco-tcp). Working with this historical corpus entails addressing both superficial and profound challenges, ranging from orthographic variation to difficulties in author attribution and challenges in field specification given the loose disciplinary boundaries of the eighteenth century.

Through this analysis, we aim to identify the changing objects of progress and how the concept was referred to throughout the century. Based on existing work in conceptual history and philosophy, we hypothesise that during the eighteenth century progress shifted from having a mostly field-specific sense to a generalised and abstract usage through processes such as metaphorisation and amelioration, while retaining its semantic core. Further, we expect to find that authors who are, in general, regarded as highly influential in this century will be some of the most salient drivers of the change. Likewise, the larger fields of science, religion, and art have been stated as areas where the concept was greatly used and therefore might be more subject to being driving forces of change.

More broadly, our study illustrates how the computational and distributional methods used in diachronic lexical semantics can provide intellectual historians with large-scale descriptive evidence that broadens the scope and complements traditional close-reading methods. In doing so, we also make an initial contribution to establishing the theoretical link between diachronic semantics and conceptual history, showing how the semantic evolution of progress reflects the intellectual and cultural dynamics of eighteenth-century Britain.


Corresponding author: Ángela María Gómez Zuluaga, KU Leuven, Department of Linguistics, Quantitative Lexicology and Variational Linguistics, Blijde-Inkomststraat 21, bus 3301, Leuven, 3000, Belgium. E-mail: angelamaria.gomezzuluaga@kuleuven.be. https://orcid.org/0009-0009-7209-4476

** KU Leuven, Department of Linguistics, Quantitative Lexicology and Variational Linguistics


References

Geeraerts, D. (1997). Diachronic prototype semantics: A contribution to historical lexicology. Oxford University Press. https://doi.org/10.1093/oso/9780198236528.001.0001 

Geeraerts, D., Speelman, D., Heylen, K., Montes, M., De Pascale, S., Franco, K., & Lang, M. (2024). Lexical Variation and Change: A Distributional Semantic Approach. Oxford University Press. https://doi.org/10.1093/oso/9780198890676.001.0001.

Koselleck, R. (2004). Futures past: On the semantics of historical time (K. Tribe, Trans.). Columbia University Press. (Original work published 1979)

Kuukkanen, J.-M. (2008). Making Sense of Conceptual Change. History and Theory, 47(3), 351–372. https://doi.org/10.1111/j.1468-2303.2008.00459.x

Marjanen, J. 2023. Quantitative Conceptual History: On Agency, Reception, and Interpretation. Contributions to the History of Concepts. 18(1). 46–67. https://doi.org/10.3167/choc.2023.180103

Nisbet, R. (2017). History of the Idea of Progress. Routledge.

Spadafora, D. (1990). The idea of progress in eighteenth-century Britain. Yale University Press.

Tolonen, M., Mäkelä, E., Ijaz, A., & Lahti, L. (2021). Corpus linguistics and Eighteenth Century Collections Online (ECCO). Research in Corpus Linguistics, 9(1), 19–34. https://doi.org/10.32714/ricl.09.01.03

Wagner, P. (2016). Progress: A reconstruction. Polity.

   

Underlying Dimensions in Semantic Change

KU Leuven

Lexical semantic change detection is a well-established field in computational linguistics and natural language processing, with several shared tasks and curated datasets (e.g., Basile et al, 2020; Kutuzov & Pivovarova, 2021; Schlechtweg et al, 2020). Modern approaches in the study of semantic change often consist of applications of vector space modelling (VSM) (Tahmasebi et al, 2021; Tahmasebi & Dubossarsky, 2023), which converts raw text into numerical vector representations then quantifies changes through clustering methods or distance measures (e.g., cosine distance) to those vectors.

 A key conceptual problem of the above shared tasks and models is that they operationalise ‘semantic change’ as change in the distribution of senses of words (between 2 time periods). Additionally, these approaches tend to suffer from the sensitivity to corpus size (Antoniak & Mimno, 2018; Sahlgren & Lenci, 2016) and the lack of interpretability (Lenci et al, 2022).

In practice, many types of change do not map neatly onto such clear-cut sense distinctions. Subtle shifts in typical collocates, prototypical structures (Geeraerts, 1997), argument-structure (Christiansen & Joseph, 2016), or conceptual frame (Traugott, 2017) may leave a coarse-grained sense inventory intact while still reflecting substantial reorganization in how a word is used and conceptualised. For instance, meeting retains the general ORGANISED GATHERING meaning, yet its usage profile develops a digitally mediated default, with new collocates like Zoom/Teams, virtual, hybrid. A strictly sense-based operationalisation therefore risks underestimating these more graded, profile-level developments, which motivates approaches that track changes in usage profiles directly rather than inferring them only through sense annotations.

We present an approach that 1) is conceptually simpler, 2) requires less data than VSM methods, 3) is more direct and theory-driven in interpreting the changes across different dimensions of words’ usages and meanings.

Given a concept/lexeme, we 1) annotate the different (theory-driven) dimensions of word usages using available tools such as automatic parsers or Large Language Models (LLMs), 2) measure change in these dimensions using divergence measures, 3) interpret the changing dimensions through divergence decomposition.

To demonstrate the feasibility of this approach, we report 3 case studies: 

In the 2 case studies with SemEval dataset (Phan-Tat et al, 2026b, 2026a), we demonstrate that it can be used for lexical semantic change detection and can outperform many VSM approaches in many cases while remaining lightweight and transparent. The models are demonstrated to capture semantically relevant signals and supports interpretable explanations of the observed changes. 

In Case Study 3, we applied the framework to three chemical concepts from late eighteenth- to early nineteenth-century chemistry and demonstrated that it successfully captures the linguistic imprint of a major paradigm shift, the emergence of oxygen chemistry. We worked with a small and highly imbalanced subset of the Royal Society Corpus 6.0 Open (Fischer et al, 2020), which is challenging for many VSM approaches. 

Because the method is lightweight and relies only on off-the-shelf AI/NLP tools, it can be applied to diverse diachronic corpora and domains. Concretely, it identifies which dimensions shift, when the shifts occur, and which lexical items contribute most to the observed divergence, thereby providing testable targets and constraints for modelling assumptions and simulations.

Acknowledgements

This work was supported by the Horizon Europe MSCA Doctoral Network project CASCADE (101119511) (www.horizoncascade.net).

References 

Antoniak, M., & Mimno, D. (2018). Evaluating the Stability of Embedding based Word Similarities. Transactions of the Association for Computational Linguistics, 6, 107–119.

Basile, P., Caputo, A., Caselli, T., Cassotti, P., & Varvara, R. (2020). DIACR-Ita @ EVALITA2020: Overview of the EVALITA2020 Diachronic Lexical Semantics (DIACR-Ita) Task. In V. Basile, D. Croce, M. Maro, & L. C. Passaro (Eds.), EVALITA Evaluation of NLP and Speech Tools for Italian - December 17th, 2020 (pp. 411–419). Torino: Accademia University Press.

Christiansen, B. J., & Joseph, B. D. (2016). On the relationship between argument structure change and semantic change. Proceedings of the Linguistic Society of America, 1, 26.

Fischer, S., Knappen, J., Menzel, K., & Teich, E. (2020). The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study.

Geeraerts, D. (1997). Diachronic Prototype Semantics: A Contribution Historical Lexicology. Oxford University Press.

Kutuzov, A., & Pivovarova, L. (2021). RuShiftEval: A Shared Task on Semantic Shift Detection for Russian. In Computational linguistics and intellectual technologies.

Lenci, A., Sahlgren, M., Jeuniaux, P., Cuba Gyllensten, A., & Miliani, M. (2022). A comparative evaluation and analysis of three generations of Distributional Semantic Models. Language Resources and Evaluation, 56(4), 1269–1313.

Phan-Tat, B., Heylen, K., Geeraerts, D., Pascale, S. D., & Speelman, D. (2026a). Reframe or remain: Unsupervised lexical semantic change detection with frame semantics.

Phan-Tat, B., Heylen, K., Geeraerts, D., Pascale, S. D., & Speelman, D. (2026b). Transparent semantic change detection with dependency-based profiles.

Sahlgren, M., & Lenci, A. (2016). The Effects of Data Size and Frequency Range on Distributional Semantic Models. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (pp. 975–980). Austin, Texas: Association for Computational Linguistics.

Schlechtweg, D., McGillivray, B., Hengchen, S., Dubossarsky, H., & Tahmasebi, N. (2020). SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection. In Proceedings of the Fourteenth Workshop on Semantic Evaluation (pp. 1–23). Barcelona (online): International Committee for Computational Linguistics.

Tahmasebi, N., Borin, L., Jatowt, A., Xu, Y., & Hengchen, S. (2021). Computational approaches to semantic change.

Tahmasebi, N., & Dubossarsky, H. (2023). Computational modeling of semantic change. ( eprint: 2304.06337)

Traugott, E. C. (2017, March). Semantic Change. Oxford University Press.

Words For Membership And Belonging: Exploring Early Conceptualizations Of Citizenship Through Vecinos In Colonial Spain

University of Sheffield

In order to understand how the discourse surrounding migration has shifted and evolved, it is necessary to identify salient and insightful terms. The Historical Thesausus of Migration and its associated taxonomy aim to be a representation for that topic within the DH space. One example of such terms is vecinos, interpreted in modern Spanish as an equivalent to neighbor in English. However, experts on the 15th and 16th century and the pre-colonial Spanish period often link this term to an early conceptualisation of citizenship within the Spanish Empire (Herzog, 2003). Despite this, there is a scarcity of literature delving into the evolution and intricacies of the word which actually illuminates the link and association between vecindad and citizenship, and even less detailing how the increase in the numbers of colonial settlers to the American continent could have played a part in its evolution. According to Perl-Rosenthal and Erman (2024) it is possible to reconstruct early conceptualisations of citizenship by analysing legislative historical records and extracting other terms associated with the word which often evoke notions of love and protection to assign allegiance and membership. 

   This case study seeks to go a step further by compiling data that spans different registers and gaining insight into the former usage of vecinos within different contexts. The data sets used for this study bring together 15th and 16th century legislation in the form of ordenanzas or ordinances made by the Spanish monarchy in an effort of maintaining and expanding their domain to the overseas territory and found in the four volumes of Recopilación de Leyes de indias. The ordinances are supplemented by two other corpora compiled from letters. Digging Into Early Colonial History or DECH (Murrieta-Flores et al., 2020)  developed by Lancaster University brings together reports, letters and first-hand accounts of the clerics and officers which oversaw the efforts of the Spanish Crown while the anthology of emigrant letters published by Otte in 1996 provides a look at the operation of the word within a colloquial genre. When explored using corpus linguistics methods such as collocates and N-grams, these three distinct data-sets can provide more well-rounded evidence. The collocational analysis of vecinos speaks to the term's aboutness through its co-occurrence with other words while the N-grams show how it was employed at a more discursive level by shedding light on lexical items that are found clustered together within a text. By using them simultaneously while applying them to each corpus individually, the result aims to yield a more substantial baseline of word usage that can be compared to the interpretations established by historians. This study also intends to tie the examination of this term back to larger topics and reflect on how the evolution of the concept of citizenship was influenced by migration in the early age of colonialism, societal structures and discursive  manifestations of belonging that were built as a result of this turning point in history. 

 

Keywords: vecinos, collocations, n-grams, citizenship, corpus, migration 

 

References


 

Herzog, T. (2003). Defining nations : immigrants and citizens in early  modern Spain and Spanish America (1st ed.). Yale University Press. https://doi.org/10.12987/9780300129830

 

Murrieta-Flores, Patricia; Jiménez-Badillo, Diego; Martins, Bruno Emanuel da Graça (2020). DECM Machine Ready Corpus. figshare. Dataset. https://doi.org/10.6084/m9.figshare.12048729.v3

 

Otte, E. (1996). Cartas privadas de emigrantes a Indias 1540-1616. Fondo de Cultura Económica de México. (Original work published 1988)

 

Perl-Rosenthal, N., & Erman, S. (2024). Inventing Birthright: The Nineteenth-Century Fabrication of jus soli and jus sanguinis. Law and History Review, 42(3), 421–448. doi:10.1017/S0738248024000221 


Recopilacion de Leyes de los Reinos de las Indias (Fifth edition.). (1841). Boix.