Session 14

Thursday 14:00 - 15:30

High Tor 3

Chair: James Chetwood

A digital leap of faith? Approaching gaps and anomalies in historical and archaeological big data

University of Vienna

The present paper addresses a recurring DH-problem: how to engage with archival and legacy (archaeological) documentation that is dispersed, unevenly recorded and described in incompatible and multilingual vocabularies. Using mortuary data as an example from the ERC-funded project RELIC, which focuses on rural Christianisation and state formation. While the example itself is specific, the theoretical and digital framework presented is applicable beyond funerary archaeology, offering a DH approach to small, heterogeneous datasets in which the goal is not only digitisation, but making category use, comparability, and evidential context inspectable at scale. 

RELIC aims to advance historical and archaeological understanding of the involvement of the general population in Christianisation and state formation, while also reflecting on disciplinary epistemic conditions in the digital age. In this contribution, we present how the data infrastructure developed for the project enables contextual comparison of controversial small datasets. Extra-normative burials are an excellent example of that, because “deviancy” is not a stable empirical class but an interpretive label whose meaning depends on comparison with its normative environment. We argue that DH infrastructures can enable such comparisons by harmonising descriptors, keeping provenance visible, and supporting cross-site exploration through interface-driven analysis. 

The project draws on a combination of published and archival archaeological data, historical documents, and spatial datasets. Our focus is on sites that include early rural churches, field cemeteries, transitional burial grounds, and ecclesiastical and secular power centres and their settlements, i.e. the social and political context of religious change between the 10th and 12th centuries. Many of these datasets are dispersed across national and disciplinary boundaries in East-Central Europe and are difficult to compare directly due to uneven recording practices and often incompatible, multilingual descriptive vocabularies.

For the project, we employ an interdisciplinary, multi-scalar digital approach. The OpenAtlas system serves as both a semantic backend and a spatial database, enabling multi-domain data integration and networked querying. It was the underlying framework for the THANADOS project (Eichert, 2021; Eichert, 2024) and the data collection has been significantly extended for RELIC. Precisely in this stage of large-scale diachronic survey, differences in granularity, documentation regimes and category uses become most apparent, highlighting what often appears on the surface as gaps and anomalies in the data. This requires the implementation of a clear protocol for the harmonization and normalization of highly individualized records while also opening a possibility for their interpretative contextualization. 

The existing dataset shows considerable regional variation in how Christianisation took shape materially. For instance, transitional cemeteries—sites where older burial customs merged with emerging Christian norms—often show spatial reconfigurations: changes in grave alignment, architectural interventions, or selective reuse. These sites provide fertile ground for interpreting religious change as embedded in social practice and funerary negotiation.

From Stone to Signal: Multimodal Retrieval as Method in Archaeological Research

Centro Studi Americanistici

Archaeological knowledge lives in many forms: arguments, diagrams, maps, images, and the traces left behind. These forms do not exist in isolation; each one shapes and is shaped by the others as meaning is made. Where Digital Humanities has often centred on preserving and indexing, new multimodal systems open up possibilities for how we move through and interpret the tangled webs of complex archives.

This paper presents a multimodal retrieval architecture developed within the context of a doctoral corpus on Olmec urbanisation, part of the international Ruta de la Obsidiana research initiative. The project explores how semantic text retrieval, CLIP-based image alignment, and grounded generation can be integrated into a structured access layer over a curated archaeological archive.

The system brings together vector-based text retrieval, cross-modal similarity, diversity-aware ranking, and page-level metadata alignment to keep images and texts in meaningful proximity. An evaluation layer (“LLM-as-a-judge”) checks for fidelity and relevance, thereby supporting traceability and transparency in response generation. By separating parametric and non-parametric memory, the archive remains open to updates and accountable to its sources. These choices shape how evidence is encountered, related to and made meaningful within the archive; they are not just technical. In practice, the system consistently retrieved page-aligned text and image pairs across languages, with groundedness scores reinforcing the traceability of each citation.

At its heart, archaeology is a discipline of attention. It is about learning to notice, to read meaning in fragments, gaps, and the ways things relate.  It trains the researcher to read traces, recognise the meaning distributed across fragments, interpret absences, and discover relationships across material. In this sense, the design of computational systems for archaeology is never neutral. It encodes assumptions about what counts as proximity, relevance, and connection. By structuring retrieval across text and image simultaneously, this project treats technological architecture as an extension of interpretative practice. The system does not replace the slow work of analysis; it reorganises the conditions under which that work becomes possible.

In this field, interpretation grows from relationships: artefact to landscape, image to inscription, fragment to structure. Retrieval is not just a technical step but a way of shaping how these relationships are traced and understood. A multimodal system that honours these connections can help researchers move through archives without losing sight of the evidence that grounds their work.

This work situates the prototype within broader Digital Humanities discussions on sustainability, epistemic responsibility, and infrastructural design. Multimodal retrieval is seen here as a way to open access while preserving the archive's complexity. The goal is not just to speed things up, but to align technology with the ways archaeological knowledge is built and shared.

If archives hold memory, retrieval systems shape how we meet it. Thoughtful design is part of how we care for heritage today.

Ecologies of the Ancient Mediterranean: Building an Evidence-Aware Spatial Graph RAG Platform for Environmental Knowledge in Texts and Landscapes (500 BCE–300 CE)

University of Zurich

Large language models can now produce fluent “environmental histories” of the ancient  world on demand. For research, that fluency is also the core risk: LLM outputs routinely  collapse heterogeneous evidentiary layers, attested statements, inferred practices, and  

interpretive framings, into a single register of “facts.” In ancient environmental history,  where genres moralize scarcity, narrate disaster through conventional tropes, or encode  administrative action in formulaic documentary language, this collapse is  methodologically corrosive. This paper presents Ecologies of the Ancient  Mediterranean, a spatially aware AI platform designed to keep evidentiary distinctions  explicit while enabling map-based exploration and retrieval across textual and  archaeological datasets. 

The platform integrates three components that are often treated separately in digital  humanities workflows: (1) place-grounded text analysis, (2) knowledge-graph modelling  of evidence, and (3) retrieval-augmented generation (RAG) with auditable citations. The  core data layer combines openly licensed corpora (TEI/EpiDoc and linked-data exports  where available), a shared gazetteer backbone (Pleiades identifiers for places and  regions), and selected archaeological/landscape datasets relevant to environmental  practice (e.g., hydraulic features, settlement patterns, land-use evidence). Texts are  segmented into citable passages (“chunks”) enriched with source-type and genre  metadata (historiography, technical/agronomic writing, documentary texts), enabling  systematic comparison of how ecological phenomena are narrated and operationalized  across genres. 

A central design principle is evidence-aware annotation. The platform encodes  distinctions between: 

• attestations (what a text explicitly says, linked to exact passages), 

• inferred events/practices (what can be reasonably derived, with provenance),  and 

• interpretive framings (e.g., divine agency, mismanagement, natural cycles),  modelled as claims about discourse rather than facts about the world. 

To stabilize extraction and support interoperability, the project develops an Ancient  Ecology Vocabulary: a controlled vocabulary and light ontology for concepts such as  “water infrastructure,” “storage practice,” “hazard event,” and “scarcity rhetoric.”  Crucially, the vocabulary is designed as a reusable semantic layer: NLP labels point to  explicit definitions and identifiers, and (where feasible) the model is designed to be 

mappable to established cultural-heritage standards rather than remaining a one-off  taxonomy. 

Methodologically, the paper details a Spatial Graph-RAG architecture that combines  vector retrieval over text chunks with graph constraints and re-ranking by place, time  window, genre, and ecological category. A scholar-facing interface supports map-driven queries (e.g., “low-Nile years and repair activity”) while surfacing the underlying  evidence trail: retrieved passages, linked places, and the annotation steps that justify  any synthesis. Evaluation is built in through small, collaboratively produced gold standard annotations, assessed at the level of vocabulary concepts (precision/recall for  categories such as “canal” vs. “aqueduct,” or “famine episode” vs. “moralized scarcity  talk”). 

The paper concludes with a case study on agricultural risk and famine in Roman Italy as  a demonstration of how spatial linking and evidence separation can reveal patterned  divergences between discourse, institutional action, and material traces—without  forcing false alignment. For the Digital Humanities Congress, the contribution is both  methodological and infrastructural: a transferable pattern for building AI-assisted  research tools that remain transparent, auditable, and sustainable for humanities  scholarship.