Session 9 — Linking Challenging Data

Friday 11:30 - 13:00

High Tor 4

Chair: Katherine Rogers

Methods for Mining Messy Real World Data: Co-reference Identification Using Fuzzy Logic

De Montfort University, Leicester

As the number and volume of online museum collections grows there is an increasing imperative to improve their discoverability by finding ways of linking records that go beyond simple keyword searching.  Keyword searches are inefficient because they are prone to errors of both omission and commission. A range of more sophisticated approaches have been developed for cross collection searching including metadata harvesting, data mining, Linked Data and Application Programming Interfaces (APIs) but these variously rely on availability of well structured, consistent and standardised data and a large corpus of text.  While some heritage institutions have successfully implemented one or more of these approaches, pioneering the way for others to follow, the majority of online collection records are not amenable to such treatments because they employ different data schemas, applied inconsistently, are often fragmentary and imprecise and non-machine readable. A further challenge is that while there are many millions of individual object records, each containing many different fields, the entries in each separate field tend to be quite small, eg. dates, person names, titles, making it difficult to apply corpus based approaches to data analysis. Converting fragmentary, messy museum records into well structured data is an unlikely prospect in the immediate future because of the costs and technical skills required.  This paper reports on the results of the FuzzyPhoto project first introduced at the Digital Humanities Congress 2012 (http://www.hrionline.ac.uk/openbook/chapter/dhc2012-croft) and describes the computational methods we have developed for co-reference identification.  We are using a constrained hierarchical clustering approach that can accommodate messy real world data. Record pair similarity is calculated using four separate similarity metrics, with each metric tuned to the specific challenges of the information held in that field, and an overall record similarity metric to combine the other four. Trial results indicate that this approach is at least as effective as expert researchers in identifying potential matches between records and considerably faster. 

Identifying the Irish in Textual Records in the Absence of Direct Evidence

King’s College London

 ANN ADAMS, alias RILEY, was indicted for a Misdemeanor. Not Guilty.

                London Jury, before Mr. Recorder.

Is Ann Adams Irish? The above is a London criminal trial transcript from the Old Bailey Proceedings and is a typical example of the type of record that early nineteenth century historians of the Irish in Britain are faced with. From this record, researchers must decide the national identity of the person mentioned therein. Unfortunately for historians, much of the best evidence for identifying the Irish in London is gone. The accents have been silenced, the witnesses who could confirm the Irishness of defendants have long since died, and in most cases no one wrote it down because it wasn’t relevant to the trial. Yet in order to study the group, we need a way to decide which records are pertinent and which are not. We need a ‘sample’.

This talk compares the effectiveness of three techniques for generating that sample by classifying defendants in criminal trials as likely Irish or likely not. These three approaches: nominal record linkage, keyword searching, and surname analysis, were able to identify 1,712 individuals out of 25,000 defendants as Irish despite a lack of direct evidence, providing a usable sample for further study. Both nominal record linkage and keyword searching proved costly approaches, offering limited results. Surname analysis proved by far the most cost effective approach, and had a high degree of accuracy, but causes historians the greatest angst.

I will discuss how I was able to test these three approaches, and how a quantitative analysis of hundreds of thousands of records has allowed me to make qualitative judgments about individuals mentioned in textual records. In the process, this approach has opened up new questions for which historians can find plausible answers about the Irish and crime in Britain, without sacrificing academic rigor.

From Napoleon Conquests to the Big Brother Sabotage: Harmonization of the Dutch Historical Censuses in the Semantic Web

VU University Amsterdam

Around the turn of the 18th century, the first integral population enumeration was held in the Netherlands during the Batavian Republic. It took over 30 years before the first official census was, by royal decree, organized and conducted in 1829, and was meant to be held from then onwards every ten years. The Dutch historical censuses are the only large scale, reliable statistical datasets available about the (demographic, social and economic) history of the Netherlands, covering an all-encompassing geographical area for over two centuries (1795–1971). Not surprisingly, the currently preserved and digitized historical censuses are the most consulted historical statistics by researchers. However, the 2 288 census tables are highly disconnected and scarcely integrated in their current form. Meaningful information is still hidden in these missing table-links, meaning that this wealth of information is not reaped to its full potential. In this paper we describe the lessons learnt in CEDAR[1], a project of the Computational Humanities Programme[2], to provide solutions to these integration problems. Our system leverages semantic technologies and Linked Data practices, which allow us to convert the census tables into a graph of fine-grained Linked Census Data. Using the distributed architecture of the Web, we interlink this graph with other online historical socioeconomic and demographic Linked Datasets.

We use the information provided by these external links to guide the harmonization process in our dataset. At the same time, we investigate which historical classifications are not online yet following Web standards, and we use our census tables (on demographic structures, housing types, occupational classes and statuses, and religious denominations) to urge the need of publishing these historical classifications on the Web. Such historical hubs could increase enormously the interoperability of other datasets. Finally, we propose a querying pipeline on the resulting harmonized census dataset to enhance the data exploration work by historians and social scientists and help answering their research questions.

[1] See http://www.cedar-project.nl/

[2] See http://www.ehumanities.nl/