Large language models can now produce fluent “environmental histories” of the ancient world on demand. For research, that fluency is also the core risk: LLM outputs routinely collapse heterogeneous evidentiary layers, attested statements, inferred practices, and
interpretive framings, into a single register of “facts.” In ancient environmental history, where genres moralize scarcity, narrate disaster through conventional tropes, or encode administrative action in formulaic documentary language, this collapse is methodologically corrosive. This paper presents Ecologies of the Ancient Mediterranean, a spatially aware AI platform designed to keep evidentiary distinctions explicit while enabling map-based exploration and retrieval across textual and archaeological datasets.
The platform integrates three components that are often treated separately in digital humanities workflows: (1) place-grounded text analysis, (2) knowledge-graph modelling of evidence, and (3) retrieval-augmented generation (RAG) with auditable citations. The core data layer combines openly licensed corpora (TEI/EpiDoc and linked-data exports where available), a shared gazetteer backbone (Pleiades identifiers for places and regions), and selected archaeological/landscape datasets relevant to environmental practice (e.g., hydraulic features, settlement patterns, land-use evidence). Texts are segmented into citable passages (“chunks”) enriched with source-type and genre metadata (historiography, technical/agronomic writing, documentary texts), enabling systematic comparison of how ecological phenomena are narrated and operationalized across genres.
A central design principle is evidence-aware annotation. The platform encodes distinctions between:
• attestations (what a text explicitly says, linked to exact passages),
• inferred events/practices (what can be reasonably derived, with provenance), and
• interpretive framings (e.g., divine agency, mismanagement, natural cycles), modelled as claims about discourse rather than facts about the world.
To stabilize extraction and support interoperability, the project develops an Ancient Ecology Vocabulary: a controlled vocabulary and light ontology for concepts such as “water infrastructure,” “storage practice,” “hazard event,” and “scarcity rhetoric.” Crucially, the vocabulary is designed as a reusable semantic layer: NLP labels point to explicit definitions and identifiers, and (where feasible) the model is designed to be
mappable to established cultural-heritage standards rather than remaining a one-off taxonomy.
Methodologically, the paper details a Spatial Graph-RAG architecture that combines vector retrieval over text chunks with graph constraints and re-ranking by place, time window, genre, and ecological category. A scholar-facing interface supports map-driven queries (e.g., “low-Nile years and repair activity”) while surfacing the underlying evidence trail: retrieved passages, linked places, and the annotation steps that justify any synthesis. Evaluation is built in through small, collaboratively produced gold standard annotations, assessed at the level of vocabulary concepts (precision/recall for categories such as “canal” vs. “aqueduct,” or “famine episode” vs. “moralized scarcity talk”).
The paper concludes with a case study on agricultural risk and famine in Roman Italy as a demonstration of how spatial linking and evidence separation can reveal patterned divergences between discourse, institutional action, and material traces—without forcing false alignment. For the Digital Humanities Congress, the contribution is both methodological and infrastructural: a transferable pattern for building AI-assisted research tools that remain transparent, auditable, and sustainable for humanities scholarship.