The “Spirit” in the Chamber: Capturing Zeitgeist in UK Parliamentary Debates

 Zeitgeist, or the “spirit of the time,” is a concept commonly used in historical and sociological research. However, its definition remains vague, domain-dependent, and its substance lacks bottom-up empirical evidence, especially from large-scale datasets. This lack of empirical grounding is a significant issue, as humanities scholars are often criticized for biases and a "cherry-picking" of evidence. Furthermore, the Zeitgeist, as a significant yet abstract force, is likely to both influence and be reflected by important semantic phenomena over time.

In this study, Zeitgeist is redefined as the cultural, political, moral, and intellectual climate of a time period, realized by a certain group of language users by means of discourse. The Zeitgeist, in this case, will be determined through proxies, which are lists of keywords and key semantic domains. This mixed-methods study draws upon recent advancements in corpus linguistics. The process begins with the determination of keywords and key semantic categories using a keyness analysis. In quantitative corpus linguistics, a keyword is a word that is statistically “characteristic of a text or a group of texts” (Culpeper & Demmen, 2015, p. 90). Following Gries (2024), this keyness will be treated as a weighted, three-dimensional measure rather than a monodimensional one, using the versatile Kullback-Leibler Divergence (KLD) (Kullback & Leibler, 1951) as the core statistical measure. Additionally, as a diachronic study, the data from the previous period is treated as the reference corpus of the past, whereas that of the period right after is the target corpus of the new time. A key innovation of this project is the incorporation of semantic tags into the pipeline, providing another dimension of information. Further scrutiny of the concordance lines, enhanced by encyclopedic knowledge, is conducted if necessary to shed light on areas that the quantitative approach fails to illuminate. The expected outcome is a list of keywords with key semantic categories for each time slice, and together, they provide the basis for tracking the changing Zeitgeist in British parliamentary debates.

The data comes from the Semantic Hansard corpus, the largest publicly available corpus which contains nearly all speeches given in the UK Parliament spanning from 1803 to 2005 and is annotated by the USAS tagger (Rayson, 2008) with semantic tags. This unique source is particularly suited for capturing Zeitgeists as it represents the significant concerns within a national legislative body, which in turn aims to represent broader public concerns. In contrast, more specialized or domain-dependent datasets often lack the longitudinal scale and broad representative scope required to provide the bottom-up empirical grounding necessary for a macroscopic cultural analysis. The preliminary analysis focuses on debates occurring during Thatcher’s, Callaghan’s, and Blair’s governments.

The study’s contributions are theoretical, methodological, and empirical. Theoretically, it refines the abstract concept of Zeitgeist into a data-backed, empirically verifiable construct, informing both conceptual history and corpus linguistics. Methodologically, the study evaluates the KLD as a versatile, bottom-up statistical tool, especially when used in conjunction with rich semantic annotation. Empirically, this research will deliver concrete, data-driven accounts of macroscopic British political Zeitgeists across governments.

References 

Culpeper, J., & Demmen, J. (2015). Keywords. In D. Biber & R. Reppen (Eds.), The Cambridge Handbook of English Corpus Linguistics (1st ed., pp. 90–105). Cambridge University Press. https://doi.org/10.1017/CBO9781139764377.006

Gries, S. T. (2024). Frequency, dispersion, association, and keyness: Revising and tupleizing corpus-linguistic measures. John Benjamins Publishing Company.

Kullback, S. & Richard A. L. (1951). On information and sufficiency. The Annals of Mathematical Statistics 22(1). 79–86.

Rayson, P. (2008). From key words to key semantic domains. International Journal of Corpus Linguistics. 13(4). 519-549. https://doi.org/10.1075/ijcl.13.4.06ray