Session 12

Thursday 11:30 - 13:00

High Tor 4

Chair: Isabella Magni

Modeling Theatrical Institutions: A Relational Digital Infrastructure for the Metastasio Theatre Archive (1827–1861)

Università degli Studi di Firenze

The study of historical theatrical archives presents significant methodological challenges, particularly when addressing heterogeneous and structurally complex collections. The Fondo Metastasio documents the institutional, administrative, and cultural activity of the Teatro Metastasio (Prato), including deliberations, correspondence, accounting records, technical documentation, and legal acts. The investigation focuses on the period between 1827 and 1860, a chronological framework identified as most suitable for constructing the historical and interpretative structure of the research. This temporal delimitation was chosen due to the breadth, consistency, and heterogeneity of the documentary materials. 

Traditional archival research emphasizes provenance, original order, and material features, all essential for historical interpretation. Yet these principles alone do not fully illuminate the relational and functional configurations that underlie the records. This paper presents the design and implementation of a relational digital infrastructure that transforms this complex corpus into a structured, queryable research environment while preserving archival integrity.

The project models the theatre as a socio-technical system. Archival documents are digitized and described, then restructured within a relational database that integrates entities such as individuals, administrative bodies, events, contracts, and physical spaces. Rather than privileging document typology alone, the system encodes institutional functions (deliberation, financial management, contractual negotiation, technical production) thereby enabling researchers to trace processes across documentary boundaries.

The resulting platform supports multidimensional analysis through cross-entity querying and visualization tools, including temporal sequences, correspondence networks, and institutional interaction graphs. These tools reveal patterns of governance, responsibility distribution, and production logistics that remain partially obscured in conventional archival consultation. The paper demonstrates how relational modeling can generate new historical insights into the operational dynamics of nineteenth-century theatrical institutions.

Methodologically, the project contributes to ongoing debates within the Digital Humanities concerning data modeling, interoperability, and sustainability. By combining established archival standards with functional modeling and linked data approaches, it proposes a scalable framework for representing complex cultural heritage collections. 

Beyond the specific case study, this research highlights the epistemological implications of digital modeling in humanities scholarship. Modeling is not a neutral act of transcription but an interpretative process that formalizes relationships and makes analytical assumptions explicit. By foregrounding transparency and reproducibility, the project situates digital infrastructure as both a research tool and a scholarly argument.

The Metastasio model thus serves as a prototype for the digital edition of theatrical archives conceived not merely as repositories of digitized documents, but as dynamic research environments. It demonstrates how computational techniques, relational databases, linked data, and visualization interfaces, can extend traditional archival methods, enabling new forms of inquiry into cultural and institutional history.

Matching Names in Composite Databases

Independent Researcher

The data given in museum collections and exhibition catalogues is an invaluable source of information for art historians. A large and growing number of museums have online collections, many contemporary exhibitions give online access to their catalogues, and scholars have created datasets through the digitisation of catalogues of historical exhibitions. The use of quantitative analysis to answer art historical questions utilising such data sources is an expanding field. A critical barrier to entry for art historical research involving data drawn from different sources where the researcher wants to look at the level of the individual artist is that of matching artists’ names across the collections or catalogues in which they appear. With large datasets or small teams of researchers it is often the case that manual matching of names in not practical.

Name matching across data sources is a generic problem faced in both academic and commercial environments. Variations in presentation or spelling means the problem is one of finding close rather than exact matches and of distinguishing variants of the same name from different names. Researchers have investigated this issue of ‘fuzzy’ name matching, developing measures of the similarity of names or complete solutions in practical applications, often using methods from machine learning.

In this paper I investigate the use of two machine learning classification algorithms, logistic regression and random forest, for matching artists’ names with three composite datasets. Two consist of the entries for paintings by male and female artists taken from the catalogues for 106 exhibitions in nineteenth-century Paris included in the Museé D’Orsay’s ‘Salons’ dataset. The third is of modern and contemporary art, constructed by the author, and taken from the metadata available in the online collections of 135 art museums in 32 countries. The features utilised in these models are a number of name similarity measures. One challenge faced in developing the models is that the problem of matching names is highly imbalanced, in that the number of name pairs in each dataset that are variant names for an artist are several orders of  magnitude smaller than the number of name pairs of different artists. There are also no pre-existent training and testing datasets.

For art historical research, name matching must identify the substantial majority of names for an artist with very few false identifications, or, in machine learning terminology, have very high levels of ‘recall’ and ‘precision’. In machine learning there is a trade-off between recall and precision, and it applies in this case. Neither logistic regression nor random forest meet the requirements. Rather, the best approach is using them in combination to identify candidate variant names that are scored manually. To complete the matching, these links are propagated and manual cross-checks are made.

I will present some Initial analyses of the name-matched datasets to illustrate their art historical value and show how they support future avenues of research. The approach developed is one scholars working with composite datasets could consider following or adapting.

From Scholarly Edition to Digital Ecosystem: Sustaining and Reusing Founding Era Primary Sources in the Digital Humanities

University of Virginia

Documentary editing projects have long been central to humanities research, producing deeply structured, peer-reviewed editions of primary sources that underpin scholarship in history, literature, political theory, and cultural studies. Yet despite their methodological rigor, these projects are often treated as static end products rather than as dynamic digital research environments. This paper presents a digital humanities case study drawn from a multi-year collaborative initiative at the University of Virginia to sustain, extend, and repurpose Founding Era and Early Republic documentary editions in anticipation of the 250th anniversary of the Declaration of Independence.

The initiative brings together multiple long-running documentary editing projects, a university press digital imprint, and shared digital publishing infrastructure to create forgingUS (http://forgingus.org/), a freely accessible public platform that reuses and recontextualizes rigorously edited primary sources. The paper examines how traditional editorial workflows—accurate transcription, textual collation, annotation, authority control, metadata creation, and indexing—are leveraged as a foundation for new digital humanities applications, including GIS-based StoryMaps, interactive timelines, correspondence network visualizations, and Linked Open Data integrations.

Rather than developing public-facing digital projects as detached interpretive layers, this work adopts an ecosystem model in which scholarly editions, technical infrastructure, and digital tools are designed to function together. The paper details how standards-based data creation (TEI/XML, controlled vocabularies, persistent identifiers, and structured metadata) enables reuse across platforms and supports interoperability with external datasets. It also addresses challenges inherent in transforming scholarly apparatus—such as footnotes, headnotes, and documentary annotations—into intellectually accessible digital experiences without sacrificing editorial transparency or historical nuance.

Particular attention is paid to best practices in sustainability, accessibility, and preservation. The project emphasizes shared infrastructure, modular workflows, and multiple-output publication strategies that allow content to circulate across print, subscription digital platforms, and open access websites. Accessibility and usability testing inform design decisions, ensuring that materials serve scholars, educators, students, and general audiences alike. The paper also considers the policy, funding, and labor implications, arguing that sustained investment in editorial infrastructure is essential for equitable, long-term digital humanities research.

By foregrounding documentary editions as reusable data and ongoing research environments, this case study contributes to broader digital humanities conversations about scale, longevity, and impact. It demonstrates how editorial scholarship can support emerging technologies—such as GIS, network analysis, and linked data—while maintaining methodological rigor. Ultimately, the paper argues that documentary editing projects, when embedded within robust digital ecosystems, offer a powerful model for sustainable, standards-driven, and publicly engaged digital humanities research