Methods for Mining Messy Real World Data: Co-reference Identification Using Fuzzy Logic

As the number and volume of online museum collections grows there is an increasing imperative to improve their discoverability by finding ways of linking records that go beyond simple keyword searching.  Keyword searches are inefficient because they are prone to errors of both omission and commission. A range of more sophisticated approaches have been developed for cross collection searching including metadata harvesting, data mining, Linked Data and Application Programming Interfaces (APIs) but these variously rely on availability of well structured, consistent and standardised data and a large corpus of text.  While some heritage institutions have successfully implemented one or more of these approaches, pioneering the way for others to follow, the majority of online collection records are not amenable to such treatments because they employ different data schemas, applied inconsistently, are often fragmentary and imprecise and non-machine readable. A further challenge is that while there are many millions of individual object records, each containing many different fields, the entries in each separate field tend to be quite small, eg. dates, person names, titles, making it difficult to apply corpus based approaches to data analysis. Converting fragmentary, messy museum records into well structured data is an unlikely prospect in the immediate future because of the costs and technical skills required.  This paper reports on the results of the FuzzyPhoto project first introduced at the Digital Humanities Congress 2012 (http://www.hrionline.ac.uk/openbook/chapter/dhc2012-croft) and describes the computational methods we have developed for co-reference identification.  We are using a constrained hierarchical clustering approach that can accommodate messy real world data. Record pair similarity is calculated using four separate similarity metrics, with each metric tuned to the specific challenges of the information held in that field, and an overall record similarity metric to combine the other four. Trial results indicate that this approach is at least as effective as expert researchers in identifying potential matches between records and considerably faster.