Identifying essay-length reprinting across eighteenth-century books and newspapers: A case study of David Hume

Large-scale text reuse detection has opened new possibilities for studying textual circulation in eighteenth-century print culture, but the reuse data itself often remains difficult to interpret. In practice, the combination of OCR noise and errors inherited from earlier BLAST-based matching produces reuse pairs that are highly fragmented, noisy, and uneven in granularity. Instead of revealing coherent acts of republication, the data frequently consists of many small and disconnected matches whose relationship to one another is unclear. This makes it difficult to distinguish substantial reprinting from incidental overlap, quotation, or other forms of textual borrowing, and it limits the historical usefulness of reuse detection at scale.

This paper presents an iterative workflow designed to recover meaningful reprint signals from such noisy reuse data and to move from pairwise matches to text-level reprint events. We begin with a pilot corpus centered on David Hume, drawing on precomputed text reuse data from Eighteenth Century Collections Online (ECCO) and extending the analysis to historical newspapers. Hume provides a focused test case for method development, but the broader objective is to build a reusable procedure for analyzing reprinting across heterogeneous eighteenth-century corpora rather than to isolate one author alone.

Our approach treats existing reuse pairs not as final evidence, but as incomplete traces of larger transmission processes. We therefore ask how fragmented matches can be reassembled into more coherent units of analysis. The workflow combines several stages: filtering noisy matches, grouping adjacent or overlapping reuse fragments, tracing their relation to likely textual origins, and aggregating them into candidate reprint structures that can be examined at the level of essays or other meaningful textual units. In this process, the crucial methodological issue is not simply whether two strings match, but whether multiple local matches collectively indicate a single act of reprinting. By shifting attention from isolated reuse pairs to structured constellations of reuse evidence, we aim to reconstruct the larger textual events obscured by noisy data.

Extending the analysis from books to newspapers is central to this study. Reprinting in eighteenth-century print culture did not occur within a single medium, and newspapers are essential for understanding how texts circulated, were excerpted, reshaped, and redistributed. At the same time, newspapers introduce additional challenges: shorter textual units, denser reuse environments, and particularly severe OCR problems. Developing a workflow that can operate across both books and newspapers therefore provides not only a technical solution to data fragmentation, but also a stronger basis for comparative historical interpretation.

The contribution of this paper is twofold. Methodologically, it offers a practical and generalizable workflow for repairing fragmented reuse data and identifying candidate reprint events at textual scale. Substantively, it creates new possibilities for studying eighteenth-century reprinting as a cross-media process rather than as a set of isolated pairwise similarities. By reconstructing larger reprint structures from noisy reuse traces, this project seeks to make computational text reuse more interpretable for humanities research and to support a deeper account of how texts moved through eighteenth-century print culture.

 

References:

 

Ryan, Y., Mahadevan, A., & Tolonen, M. (2023, December 22). A comparative text similarity analysis of the works of Bernard Mandeville. Digital Enlightenment Studies, 1(1), 28–58. https://doi.org/10.61147/des.6
 

Gale. (n.d.). Eighteenth Century Collections Online (ECCO). Gale Primary Sources. Retrieved November 3, 2025, from https://www.gale.com/primary-sources/eighteenth-century-collections-online
 

Gale. (n.d.). Seventeenth and Eighteenth Century Burney Newspapers Collection. Gale Primary Sources. Retrieved November 3, 2025, from https://www.gale.com/primary-sources/seventeenth-and-eighteenth-century-burney-newspapers-collection
 

Gale (with the Bodleian Library). (n.d.). Seventeenth and Eighteenth Century Nichols Newspapers Collection. Gale Primary Sources. Retrieved November 3, 2025, from https://gale.com/cn/product-catalog/primary-sources/seventeenth-and-eighteenth-century-nichols-newspaper-collection

 

Rosson, D., Mäkelä, E., Vaara, V., Mahadevan, A., Ryan, Y., & Tolonen, M. (2023, April 17). Reception Reader: Exploring text reuse in early modern British publications. Journal of Open Humanities Data, 9, 5. https://doi.org/10.5334/johd.101