Session 8

Thursday 09:30 - 11:00

High Tor 3

Chair: Sophie Whittle

Evaluating Lexical Extraction in Hiberno-English: Rule-Based and LLM- Assisted Approaches

University College Cork

This paper presents and evaluates two computational approaches for parsing A Dictionary of  Hiberno-English (Terence Patrick Dolan)(Dolan, 2006) into a structured, machine-readable  lexical resource. Although the dictionary exhibits a high degree of editorial consistency, its  entries remain computationally challenging to structure (Atkins & Rundell, 2008) . Lexical  information is distributed across semi-regular templates that are frequently disrupted by  redirects, variant spellings, homonym numbering, multi-part grammatical labels, numbered  senses, phrase entries, and entries with missing or atypical pronunciation markers. Further  complexity arises because region codes, informant identifiers, editorial asides, etymologies,  examples, and cross-references are often embedded within running prose or parenthetical  material rather than isolated in stable positions. As a result, extraction requires field recognition  and careful boundary resolution, normalization, and disambiguation (Ide & Véronis, 1995). 

To address these challenges, we implement and compare two complementary approaches: a  rule-based parser and a large language model (LLM)-assisted extraction pipeline(Brown et al.,  2020; Jurafsky & Martin, 2023). The rule-based system uses deterministic segmentation and  pattern recognition to identify core lexical components, including headwords, pronunciation,  variants, part-of-speech labels, definitions, etymologies, examples, and cross-references, while  explicitly handling recurrent irregularities such as redirect entries, variant forms, numbered  senses, compound grammatical labels, and embedded metadata. The generative pipeline, by  contrast, employs schema-driven prompting, encoding the target data model within the prompt  itself targeted at field-consistent outputs and reducing ambiguity in the model’s interpretation  of semi-structured inputs and edge cases(Ji et al., 2023). 

To enable systematic comparison, we develop an evaluation framework that captures multiple  dimensions of extraction quality, including (1) field-level presence and absence, (2) exact string  agreement between methods, (3) token-level similarity using both ordered and unordered  matching, and (4) source coverage through token- and character-level containment checks targeted at detection of overgeneration and omissions. These measures function as lightweight proxies for identifying hallucination-like behaviour in generative outputs, without requiring  exhaustive manual annotation. 

Our evaluation reveals distinct performance profiles across the two approaches. The rule-based  system achieves high precision and strong source fidelity, producing outputs that remain  consistently traceable to the input text, but is sensitive to formatting variation and requires  substantial manual rule engineering. By contrast, the LLM-based system demonstrates greater  robustness to irregular structures and achieves higher recall on complex entries but exhibits  variability in output length and occasional divergence from the source text. These behaviours  are reflected in overgeneration patterns and reduced containment scores, highlighting trade offs between flexibility and control. 

By systematically comparing these approaches, we show that LLM-based extraction can  produce usable lexical data when guided by schema-driven prompting and evaluated through  explicit, field-aware metrics. Rather than treating generative outputs as opaque, we  demonstrate how they can be assessed through multi-dimensional measures, including field 

level validation, token similarity, and source coverage. In the wake of recent LLM-driven  approaches, our findings suggest that some aspects of lexical structuring can be achieved with  reduced reliance on extensive rule engineering, opening new possibilities for working with  semi-structured historical resources. 

References 

Atkins, B. T. S., & Rundell, M. (2008). The Oxford Guide to Practical Lexicography. Oxford  University Press. 

Brown, T. B., Mann, B., Ryder, N., & al, et. (2020). Language Models are Few-Shot  Learners. Advances in Neural Information Processing Systems (NeurIPS). Dolan, T. P. (2006). A Dictionary of Hiberno-English: The Irish Use of English (2nd ed.).  Gill & Macmillan. 

Ide, N., & Véronis, J. (1995). Encoding Dictionaries. In Text Encoding Initiative:  Background and Context. Kluwer Academic Publishers.

Ji, Z., Lee, N., Frieske, R., & others. (2023). Survey of Hallucination in Natural Language  Generation. ACM Computing Surveys

Jurafsky, D., & Martin, J. H. (2023). Speech and Language Processing.

Addressing Pragmatics: Linguistic Signals of Literary Canonicity

University College Cork

  Other Authors: Yuri Bizzoni3, Stefania Degaetano-Ortlieb2, James O'Sullivan1

1 University College Cork, Ireland

2 Saarland University, Germany

3 Aarhus University, Denmark

 

Canonicity is not static; it is an exclusionary process in a given context of reading, where canon status is not a property of the work itself but of its transmission in relation to other works regarded as the repository of cultural values (Guillory, 1987). Current approaches use data-driven methods to model canonicity. Tolonen et al. (2021) used hierarchical topics to analyse how the interaction between topic development and publication-related factors explained a work’s canonical status. Similarly, Lassche et al. (2025) used LLMs to build semantic embeddings for the books and analyse how the interplay between text-intrinsic and text-extrinsic factors distinguished canonical from non-canonical works. However, improving time-awareness and interpretability to explain canonicity as an adaptive phenomenon in literary texts remains a challenge.

 

One established avenue is to analyse linguistic markers of canonicity (e.g. what authors write about and how they write about it) alongside socio-cultural factors over time (Bizzoni et al., 2025). We aim to investigate what distinctive features characterise works likely to be canonised, drawing from a corpus consisting of novels and short story collections between 1700 and 2024, written by Irish and British authors. We determined canonical status by cross-referencing each text in our corpora with multiple established markers of literary recognition, such as books published by Penguin Classics and Oxford World’s Classics. Works appearing in any of these sources were labelled canonical, while those absent were labelled non-canonical. Considering the established stylistic features, we computed average sentence lengths. For syntactic profiling, we used spaCy to perform part-of-speech tagging (Honnibal & Montani, 2017). We assessed readability using Textstat’s Flesch-Kincaid Grade Level, providing an estimate of syntactic and lexical difficulty (Bansal, 2018). Finally, we measured lexical diversity using the type-token ratio, offering a measure of vocabulary range and variation across canonical and non-canonical texts.

 

So far, we found canonical texts: (a) use longer sentences and have greater sentence length deviation, (b) use slightly more nouns, prepositions, auxiliary verbs, and pronouns, and (c) are generally less readable. In contrast, non-canonical texts use slightly more determiners, proper nouns, and a more diverse vocabulary. Notably, these results align with a similar study characterising canonicity (cf. Bizzoni et al., 2025). If these features function as indicators of stylistic stability, we expect canonical works to remain consistent over time, while non-canonical works will show greater divergence as language use evolves. To this end, we will implement information-theoretic metrics to detect lexico-grammatical change (Teich et al., 2021) and analyse to what extent this overlaps with reported socio-cultural factors.

 

We will present ongoing work on modelling canonicity, comparing the formation of the Irish and British literary canons as a case study. Rather than assuming canonical status reflects intrinsic quality, these results offer quantitative evidence that canonicity is associated with stylistic complexity, showing how critical norms privilege certain linguistic signals in the construction of literary canons, and suggesting a path toward formulating a quantitative definition for canonicity. Overall, our research contributes to improving interpretability in models explaining canon formation and long-term literary survival in the Digital Humanities.


 

References

Bansal, S. (2018). textstat: A Python library for calculating statistics from text. GitHub. https://github.com/shivam5992/textstat.

 

Bizzoni, Y., Feldkamp, P., & Nielbo, K. (2025, May). Effects of Publicity and Complexity in Reader Polarization. In Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities (pp. 138-150).

 

Guillory, J. (1987). Canonical and Non-Canonical: A Critique of the Current Debate. ELH, 54(3), 483–527. https://doi.org/10.2307/2873219.

 

Honnibal, M., & Montani, I. (2017). spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing.

 

Lassche, A., Feldkamp, P., Bizzoni, Y., Baunvig, K., & Nielbo, K. (2025, May). Why Novels (Don’t) Break Through: Dynamics of Canonicity in the Danish Modern Breakthrough (1870-1900). In Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2025) (pp. 278-290).

 

Teich, E., Fankhauser, P., Degaetano-Ortlieb, S., & Bizzoni, Y. (2021). Less is more/more diverse: on the communicative utility of linguistic conventionalization. Frontiers in Communication, 5, 620275.

 

Tolonen, M., Hill, M.J., Ijaz, A.Z., Vaara, V., Lahti, L. (2021). Examining the Early Modern Canon: The English Short Title Catalogue and Large-Scale Patterns of Cultural Production. In: Baird, I. (eds) Data Visualization in Enlightenment Literature and Culture . Palgrave Macmillan, Cham. https://doi.org/10.1007/978-3-030-54913-8_3.  

Queer Negativity Meets Vector Algebra: Using Word Embeddings to Explore the Abortion Metaphor in Literary Fiction

University of Edinburgh

As the sociologist Laura Nelson observes, machine learning and research into culture are scholarly modes of inquiry which have developed in largely orthogonal ways, but which are epistemologically aligned. As such, integrating these approaches in ways attentive to this alignment has the power to more fully realise the potential of computational literary analysis (2). Moreover, this kind of cross-disciplinary work does not freeze or limit epistemological frameworks but opens them up in productive ways, allowing researchers to, as Clayton Childress puts it, move our methods beyond triangulation and confirmatory to mechanisms for “propulsive facilitation” ranging “across methods and data sources in which epistemic standpoints change as we go” (980). These calls from the social sciences resonate with those from within computational literary studies for scholars to do what Richard Jean So and Edwin Roland describe as “the hard work of developing a critical version of distant reading” (72) capable of rising to the challenge of analysing the dynamic and contingent terms through which human identities, subjectivities and bodies are represented and categorised, in ways which are sensitive both to historical specificity and to the workings of literary form and genre. 

 

Situating itself within this line of digital humanities work, this paper takes up the word abortion, a term whose meanings and associations vary considerably across domains and across historical and cultural contexts, and considers the utility of machine learning for understanding the work it does, and the associations it carries, in different discursive environments. Emerging out of a dissatisfaction with the connotations of failure and malfunction that accrete around the term abortion, it brings together perspectives from queer theories of negativity (eg. Halberstam) which provide a framework for interrogating the conceptual underpinnings of those accretions, and work from literary history exploring ‘the abortion metaphor’ as it is articulated and fleshed out in novels and other writings of the early twentieth century, together with word embeddings, a technique from natural language processing in which words in a corpus are mapped to low-dimensional vectors and the relationships between these vectors are then used to reveal latent semantic relationships between words (Antoniak and Mimno 107). We argue that applying machine learning methods to the term abortion presents two key methodological innovations. First, they offer a way to grasp the diversity of contexts in which the term has appeared, from embryology to literary aesthetics, and suggest how these illuminate different aspects of failure or non-completion. Second, using vector algebra to subtract negative associations from the term allows us to explore the discursive terrain abortion occupies when those associations are removed. If word embeddings are an effective way of capturing the manifold contexts in which a word appears, abortion presents itself as a particularly intriguing example for literary scholars and those in the medical humanities to interrogate, inflected as it is with highly positive connotations in some contexts and highly negative ones in others. 

 

Works cited

 

Antoniak, Maria, and David Mimno. ‘Evaluating the Stability of Embedding-Based Word Similarities’. Transactions of the Association for Computational Linguistics, vol. 6, no. 0, Feb. 2018, pp. 107–19.

Childress, Clayton. ‘Bringing Computation into Cultural Theory: Four Good Reasons (and One Bad One)’. New Literary History, vol. 54, no. 1, 2022, pp. 975–83.

Halberstam, J. The Queer Art of Failure. Duke University Press, 2011.

Nelson, Laura K. ‘Leveraging the Alignment Between Machine Learning and Intersectionality: Using Word Embeddings to Measure Intersectional Experiences of the Nineteenth Century U.S. South’. Poetics, vol. 88, Oct. 2021, p. 101539. 

So, Richard Jean, and Edwin Roland. ‘Race and Distant Reading’. PMLA, vol. 135, no. 1, Jan. 2020, pp. 59–73.