Distributed cataloguing, interoperability and reuse within and beyond the infrastructural setting of the Text+ Registry

Scholarly resources are typically developed within project-specific contexts, creating a fragmented ecosystem where results remain isolated in institutional silos. This not only impedes the visibility of resources, but also limits the potential for collaborative enhancement, cross-referencing, and the sustainable stewardship of digital scholarly outputs. The Text+ Registry [1], developed within Germany's National Research Data Infrastructure [2], shifts the paradigm from centralized control to distributed curation, transforming the way scholarly communities engage with resource metadata across institutional and disciplinary boundaries.

Traditional cataloguing approaches face a critical dilemma: either enforce rigid standardization that fails to capture domain-specific nuances or accept incompatible local solutions that hinder interoperability. The Registry's architectural response to this dilemma is a metamodeling framework combined with provenance-aware information layering. Rather than flattening diverse metadata into a lowest common denominator, the system maintains multiple autonomous layers that preserve the integrity and context of each data source while enabling their synthetic combination. 

The layering system supports automated aggregation and expert curation. Manual enrichments become additional layers that enhance collective knowledge without overwriting original contributions. This reframes interoperability as post-hoc composition that respects heterogeneity, rather than prior agreement on a unified schema. The technical implementation uses a model-driven design approach, in which formal data models tailored to each data domain generate application components that are independent of specific technologies. YAML-based model definitions enable schemas to evolve without requiring a redesign of the system. Shared DataCite mappings facilitate interoperability and cross-domain queries, while preserving the richness of domain-specific metadata. Furthermore, the integration of authority files such as the Integrated Authority File [3] facilitates semantic links between resources, actors, and institutions.

This distributed curation model addresses sustainability challenges inherent in community-maintained infrastructure. Rather than requiring resources to abandon existing cataloguing practices or migrate to centralized platforms, the Registry enables participation through low-barrier mechanisms: automated harvesting from existing APIs, form-based manual submissions, and Edit-a-thon events for collaborative enhancement. Version control and transparent attribution ensure all contributions remain traceable and reversible.

Beyond its role within Text+, the metamodeling and layering capabilities have proven transferable. They have facilitated the integration of Monumenta Germaniae Historica (MGH) into the Text+ resource landscape and now underpin cataloguing systems in two further initiatives: Research data management infrastructures in Germany and the long-term project Modern India in German Archives (MIDA). These reuse cases demonstrate that the Registry’s technical framework is not bound to a single infrastructural context.

Having presented the architectural blueprint of the registry at the 2024 DHC, we now report on developments and insights gained over the past two years. The paper examines how this framework combines technological components, distributed curation and community engagement to operationalize FAIR principles in practice, moving toward concrete workflows that incentivize data sharing and collaborative stewardship. By analysing specific integration cases—from the technical challenges of MGH-to-registry transformation to the organizational dynamics of cross-institutional curation—we demonstrate how infrastructure design choices shape possibilities for scholarly collaboration.

 


1. https://registry.text-plus.org 

2. Nationale Forschungsdateninfrastruktur (NFDI), https://www.nfdi.de/?lang=en 

3. Gemeinsame Normdatei (GND), https://www.dnb.de/EN/Professionell/Standardisierung/GND/gnd_node.html