Addressing Pragmatics: Linguistic Signals of Literary Canonicity

  Other Authors: Yuri Bizzoni3, Stefania Degaetano-Ortlieb2, James O'Sullivan1

1 University College Cork, Ireland

2 Saarland University, Germany

3 Aarhus University, Denmark

 

Canonicity is not static; it is an exclusionary process in a given context of reading, where canon status is not a property of the work itself but of its transmission in relation to other works regarded as the repository of cultural values (Guillory, 1987). Current approaches use data-driven methods to model canonicity. Tolonen et al. (2021) used hierarchical topics to analyse how the interaction between topic development and publication-related factors explained a work’s canonical status. Similarly, Lassche et al. (2025) used LLMs to build semantic embeddings for the books and analyse how the interplay between text-intrinsic and text-extrinsic factors distinguished canonical from non-canonical works. However, improving time-awareness and interpretability to explain canonicity as an adaptive phenomenon in literary texts remains a challenge.

 

One established avenue is to analyse linguistic markers of canonicity (e.g. what authors write about and how they write about it) alongside socio-cultural factors over time (Bizzoni et al., 2025). We aim to investigate what distinctive features characterise works likely to be canonised, drawing from a corpus consisting of novels and short story collections between 1700 and 2024, written by Irish and British authors. We determined canonical status by cross-referencing each text in our corpora with multiple established markers of literary recognition, such as books published by Penguin Classics and Oxford World’s Classics. Works appearing in any of these sources were labelled canonical, while those absent were labelled non-canonical. Considering the established stylistic features, we computed average sentence lengths. For syntactic profiling, we used spaCy to perform part-of-speech tagging (Honnibal & Montani, 2017). We assessed readability using Textstat’s Flesch-Kincaid Grade Level, providing an estimate of syntactic and lexical difficulty (Bansal, 2018). Finally, we measured lexical diversity using the type-token ratio, offering a measure of vocabulary range and variation across canonical and non-canonical texts.

 

So far, we found canonical texts: (a) use longer sentences and have greater sentence length deviation, (b) use slightly more nouns, prepositions, auxiliary verbs, and pronouns, and (c) are generally less readable. In contrast, non-canonical texts use slightly more determiners, proper nouns, and a more diverse vocabulary. Notably, these results align with a similar study characterising canonicity (cf. Bizzoni et al., 2025). If these features function as indicators of stylistic stability, we expect canonical works to remain consistent over time, while non-canonical works will show greater divergence as language use evolves. To this end, we will implement information-theoretic metrics to detect lexico-grammatical change (Teich et al., 2021) and analyse to what extent this overlaps with reported socio-cultural factors.

 

We will present ongoing work on modelling canonicity, comparing the formation of the Irish and British literary canons as a case study. Rather than assuming canonical status reflects intrinsic quality, these results offer quantitative evidence that canonicity is associated with stylistic complexity, showing how critical norms privilege certain linguistic signals in the construction of literary canons, and suggesting a path toward formulating a quantitative definition for canonicity. Overall, our research contributes to improving interpretability in models explaining canon formation and long-term literary survival in the Digital Humanities.


 

References

Bansal, S. (2018). textstat: A Python library for calculating statistics from text. GitHub. https://github.com/shivam5992/textstat.

 

Bizzoni, Y., Feldkamp, P., & Nielbo, K. (2025, May). Effects of Publicity and Complexity in Reader Polarization. In Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities (pp. 138-150).

 

Guillory, J. (1987). Canonical and Non-Canonical: A Critique of the Current Debate. ELH, 54(3), 483–527. https://doi.org/10.2307/2873219.

 

Honnibal, M., & Montani, I. (2017). spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing.

 

Lassche, A., Feldkamp, P., Bizzoni, Y., Baunvig, K., & Nielbo, K. (2025, May). Why Novels (Don’t) Break Through: Dynamics of Canonicity in the Danish Modern Breakthrough (1870-1900). In Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2025) (pp. 278-290).

 

Teich, E., Fankhauser, P., Degaetano-Ortlieb, S., & Bizzoni, Y. (2021). Less is more/more diverse: on the communicative utility of linguistic conventionalization. Frontiers in Communication, 5, 620275.

 

Tolonen, M., Hill, M.J., Ijaz, A.Z., Vaara, V., Lahti, L. (2021). Examining the Early Modern Canon: The English Short Title Catalogue and Large-Scale Patterns of Cultural Production. In: Baird, I. (eds) Data Visualization in Enlightenment Literature and Culture . Palgrave Macmillan, Cham. https://doi.org/10.1007/978-3-030-54913-8_3.