Session 10Thursday 11:30 - 13:00High Tor 2Chair: Susan Fitzmaurice (Professor) |
|---|
Infinite love and saving faith: the conceptual exploration of English religious vocabulary through digital resourcesUniversity of GlasgowLucy Hutchinson (1620-1681) was until comparatively recently best-known for her biography of her husband John (d.1664), an officer in the New Model Army and one of the regicides who had signed the warrant for the execution of Charles I. The ongoing publication of her writings (Clarke et al 2019) has, however, made her much more famous, revealing that she was in her own right a profoundly learned individual. She wrote widely across several genres in addition to biography, with outputs including in addition not only a substantial body of poetry but also a translation of Lucretius’s De rerum natura, and important religious treatises such as On the Principles of the Christian Religion and On Theology.
Hutchinson was in all her writing highly sensitive to linguistic nuance, and the proposed case-study, part of a large-scale survey of religious discourse funded by the Leverhulme Trust to be published in late 2026 (Smith forthcoming), locates and contextualises her usage diachronically, focusing, because of time-constraints, on two English phrases of importance in her repertoire: infinite love and saving faith. Neither phrase is scriptural, but both over time have developed special theological significance. This process of cultural mapping or conceptual history – a contribution to the recuperation of what has been called historical theolinguistics – draws on major corpora such as EEBO-TCP, CLMET and COCA, as well as specially curated resources, and harnesses powerful linked digital tools, notably Linguistic DNA, Semantic EEBO and the Historical Thesaurus of English.
References
Clarke, Elizabeth, David Norbrook and Jane Stevenson (eds) 2019. The Works of Lucy Hutchinson: Theological Writings and Translations (Oxford: University Press) Smith, Jeremy J. forthcoming. Lexicons of English Religion 1380-1850 (Cambridge: University Press) |
Beyond Lexical Reuse: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of LockeUniversity of HelsinkiThe digitization of vast historical archives, such as the Eighteenth Century Collections Online (ECCO), offers unprecedented opportunities to trace the dissemination of ideas. However, traditional computational approaches to reception history rely heavily on lexical text reuse detection, which faces the challenges of implicit reference, significant OCR errors, diachronic semantic change, and the conceptual complexity of intellectual language. This challenge is particularly acute when tracing John Locke, whose complex arguments often appear in implicit paraphrases that evade lexical retrieval. To address this, we investigate semantic search as a complementary tool to close reading, evaluating its utility for tracking the reception of complex ideas. This study presents the first large-scale validation of semantic retrieval on ECCO, pioneering a paradigm beyond lexical reuse. Our methodology utilizes the entire ECCO corpus (32 million pages), enriched with ESTC metadata. We segmented the corpus into 100-token chunks and encoded them using a Sentence-BERT model, constructing a FAISS index for efficient dense vector retrieval at scale. We implemented a multi-stage evaluation, beginning with the selection of 20 candidate quotes from a pool of 1,000 frequent reuses. For each candidate, we retrieved the top 200,000 semantically similar chunks, filtering out lexical reuses to isolate novel matches. Following a stratified manual annotation of 50 hits per quote, we selected two high-performing quotes for deeper analysis, considering both high retrieval quality and conceptual richness. We annotated a denser sample of the top 200 hits each, yielding two curated datasets of approximately 400 validated instances in total.
Our initial analysis reveals that while the relationship between similarity predictions and semantic categories is noisy, the signal-to-noise ratio remains favorable. Crucially, in 15 of 20 sampled quotes, the top-ranked hits were predominantly valid matches. In the further macro-level analysis of 400 curated instances, we found that the volume of works containing semantic reuse significantly exceeds that of lexical reuse. The distribution of works varies notably by genre and time; for instance, one quote shows a distinct peak in the 1740s within the science genre, linked to the rise of Newtonianism.
This study offers both a methodological proof-of-concept and a rigorous empirical resource. We demonstrate that a workflow combining semantic search with efficient annotation can surpass manual retrieval to discover valuable non-lexical hits. Future work will leverage these datasets as a gold standard to develop LLM-assisted annotation pipelines, allowing us to scale from single quotes to broad intellectual traditions. |
Identifying essay-length reprinting across eighteenth-century books and newspapers: A case study of David HumeUniversity of HelsinkiLarge-scale text reuse detection has opened new possibilities for studying textual circulation in eighteenth-century print culture, but the reuse data itself often remains difficult to interpret. In practice, the combination of OCR noise and errors inherited from earlier BLAST-based matching produces reuse pairs that are highly fragmented, noisy, and uneven in granularity. Instead of revealing coherent acts of republication, the data frequently consists of many small and disconnected matches whose relationship to one another is unclear. This makes it difficult to distinguish substantial reprinting from incidental overlap, quotation, or other forms of textual borrowing, and it limits the historical usefulness of reuse detection at scale. This paper presents an iterative workflow designed to recover meaningful reprint signals from such noisy reuse data and to move from pairwise matches to text-level reprint events. We begin with a pilot corpus centered on David Hume, drawing on precomputed text reuse data from Eighteenth Century Collections Online (ECCO) and extending the analysis to historical newspapers. Hume provides a focused test case for method development, but the broader objective is to build a reusable procedure for analyzing reprinting across heterogeneous eighteenth-century corpora rather than to isolate one author alone. Our approach treats existing reuse pairs not as final evidence, but as incomplete traces of larger transmission processes. We therefore ask how fragmented matches can be reassembled into more coherent units of analysis. The workflow combines several stages: filtering noisy matches, grouping adjacent or overlapping reuse fragments, tracing their relation to likely textual origins, and aggregating them into candidate reprint structures that can be examined at the level of essays or other meaningful textual units. In this process, the crucial methodological issue is not simply whether two strings match, but whether multiple local matches collectively indicate a single act of reprinting. By shifting attention from isolated reuse pairs to structured constellations of reuse evidence, we aim to reconstruct the larger textual events obscured by noisy data. Extending the analysis from books to newspapers is central to this study. Reprinting in eighteenth-century print culture did not occur within a single medium, and newspapers are essential for understanding how texts circulated, were excerpted, reshaped, and redistributed. At the same time, newspapers introduce additional challenges: shorter textual units, denser reuse environments, and particularly severe OCR problems. Developing a workflow that can operate across both books and newspapers therefore provides not only a technical solution to data fragmentation, but also a stronger basis for comparative historical interpretation. The contribution of this paper is twofold. Methodologically, it offers a practical and generalizable workflow for repairing fragmented reuse data and identifying candidate reprint events at textual scale. Substantively, it creates new possibilities for studying eighteenth-century reprinting as a cross-media process rather than as a set of isolated pairwise similarities. By reconstructing larger reprint structures from noisy reuse traces, this project seeks to make computational text reuse more interpretable for humanities research and to support a deeper account of how texts moved through eighteenth-century print culture.
References:
Ryan, Y., Mahadevan, A., & Tolonen, M. (2023, December 22). A comparative text similarity analysis of the works of Bernard Mandeville. Digital Enlightenment Studies, 1(1), 28–58. https://doi.org/10.61147/des.6 Gale. (n.d.). Eighteenth Century Collections Online (ECCO). Gale Primary Sources. Retrieved November 3, 2025, from https://www.gale.com/primary-sources/eighteenth-century-collections-online Gale. (n.d.). Seventeenth and Eighteenth Century Burney Newspapers Collection. Gale Primary Sources. Retrieved November 3, 2025, from https://www.gale.com/primary-sources/seventeenth-and-eighteenth-century-burney-newspapers-collection Gale (with the Bodleian Library). (n.d.). Seventeenth and Eighteenth Century Nichols Newspapers Collection. Gale Primary Sources. Retrieved November 3, 2025, from https://gale.com/cn/product-catalog/primary-sources/seventeenth-and-eighteenth-century-nichols-newspaper-collection
Rosson, D., Mäkelä, E., Vaara, V., Mahadevan, A., Ryan, Y., & Tolonen, M. (2023, April 17). Reception Reader: Exploring text reuse in early modern British publications. Journal of Open Humanities Data, 9, 5. https://doi.org/10.5334/johd.101
|