Thursday, 3 July 2014

RDM Technical Infrastructure Review - Architecture

2.     RDM Technical Infrastructure Architecture

2.1.   Technical infrastructure components

Generally, the infrastructure architecture cases examined below have been designed around functional requirements derived from researcher workflows. Some innovative infrastructure platforms have been architected to manage research data throughout the lifecycle, but most infrastructure projects have had to take into account current systems and have engineered modifications to facilitate interoperability (Hitchcock, 2012).

2.1.1.      Major functional components of the RDM infrastructure:

·         Metadata Capture [4.7] - Metadata capture (data cataloguing) may be accomplished simply by providing an interface for researchers to fill out online forms. It is best for this process to be automated where possible to reduce the amount of manual annotation required of researchers. As well as reducing ‘double-keying’, which is frustrating for researchers, the number of errors introduced (inevitable through manual input) is reduced. Automatic metadata capture, concurrent with data capture, may be facilitated by using appropriate instruments and equipment and save data to the laboratory, departmental or facility file store or the institutional network. Electronic lab books, electron microscopes and other imaging instruments, genetic sequencing and analysis instruments may feed data to a project based Laboratory Information Management Systems (LIMS) [4.7.1].

·         Active Research Data Management [4.4] - Active research data needs to be accessed rapidly, may require large computational resources and require stringent security and access arrangements. Many institutions have developed collaborative computing systems, such as Virtual Research Environments (VREs) or LIMS, to accommodate these needs. Active data management is comprised of two functional components: a filestore and a data registry (or metadata store or asset registry). In some cases these components will be integrated into a single system, in other cases, the metadata may be handled by the CRIS.

·         Research Data Repository [4.2] - Archive data is possibly best managed in a discipline based repository or data centre, whilst the Institutional repository is the ‘repository of last resort’, as previously discussed. The institutional data repository (or the institutional repository, if it has been modified to accommodate data) is an appropriate home for datasets for which there is no discipline based repository or data centre, or for temporary storage before being submitted to a data centre. The Research Data Repository provides a catalogue for all published data and possibly file storage for some published data. The catalogue and archive functions of the repository may be separated.

·         Research Data Catalogue [4.5] - holds the metadata records of published research data. The data themselves may be held in a discipline-based data repository outside the institution or in an institutional data archive. The selection of the underlying metadata schema is fundamental and consideration must be given to the schema used by the proposed National Data Registry[i]. Many institutions favour the Datacite metadata schema[ii], subscription to which provides the means to mint DOIs and assurance of a standard level of preservation.

·         Research Data Archive [4.3] - preserves data not, or not yet, submitted to discipline-based data repositories. The associated metadata records are held in the research data catalogue.

·         Current Research Information System (CRIS) [4.6] - manages the metadata associated with researcher identity, project information, research costing, grant applications and awards. CRISs hold details of researchers’ published outputs with associated citation metrics.

All components of this technical infrastructure need to be interoperable. This is achieved through adherence to data and metadata formats and standards allowing data and metadata exchange between interoperable systems.
Many implementations involve an overlap in these functional components; for example the CRIS may provide aspects of active data management, providing a research data registry (inward facing); Laboratory instruments may be part of a LIMS, so that automated data capture is an integral part of the collaborative active data management system; Data storage and data catalogue (outward facing) may be separate systems or may be combined in a data repository. The storage / archive function of the data repository may be achieved using an external archive service, but access and deposit managed seamlessly through the repository platform (see figure 3).

Figure 3. RDM technical infrastructure data flow 

2.1.2. Data grids and Micro-services

The data grid is a form of RDMI architecture in which middleware applications allow researchers to manage data across grid infrastructures. Grid computing involves a distributed infrastructure served by interoperable software services, ‘middleware’, allowing resource sharing; the resulting ‘Grid’ may be considered a ‘virtual organisation’ (Foster et al. 2003). Data grids permit the sharing of computational resources, storage resources, network resources, code repositories and catalogues. Access to Grid resources are controlled by a Resource management system, or Storage resource broker (SRB). Several middleware toolkits are available, including open source options Globus [4.10.1] and MyGrid [4.10.2].

The micro-services architecture approach considers the repository as a ‘set of services’ rather than a ‘place’ (Abrams et al. 2009). Each function in a workflow is embodied in a self-contained micro-service, which is joined with other micro-services in a ‘pipeline’ to produce complex processes. The micro-services approach has been developed by the California Curation Centre[iii] and put into practice at the University of California Digital Library (CDL) Merritt repository [6.3.5]. At the University of Oxford, micro-services are built on an underlying Fedora repository platform, creating the Databank repository system [4.2.4], used as the platform for Oxford Databank [6.1.12].

The iRODS [4.1.3] software system allows the management of a distributed workflow through the chaining of micro-services. iRODS software is termed ‘adaptive middleware’ and allows for a more flexible customisation of data management functions than can be achieved using a SRB system. These functions or micro-services are coded as ‘rules’, which may be compiled together to produce larger macro-level functionality. 

Hydra [4.2.5] is multi-purpose repository framework based on a micro-services architecture. The main components are a Fedora repository platform [4.2.3], SOLR indexing software [4.10.3], Blacklight discovery interface [4.5.5] and the Hydra plugin, a ‘Ruby on rails’ library, which facilitates workflow in digital object management (Awre, 2012). Hydra has been implemented at the University of Hull [6.1.9] and the University of Virginia [6.3.10] with provisions made for curating research datasets. At Hull micro-systems implement workflows that allow deposit of materials via the CRIS, Converis [4.6.3], the Sakai VLE [4.4.2] and Sharepoint [4.4.3].



2.2.   Functional requirements

The functional requirements of the RDM Infrastructure may be derived from analysis of stakeholder activities, particularly researcher workflows. Many of the JISC RDMI projects have carried out data audits and investigated researcher workflows and use case scenarios in order to specify infrastructure requirements. The following list is derived from the findings of several of the JISC RDMI projects: ADMIRe (Sero Consulting, 2012; Parsons and Berry, 2012), CKAN for RDM (Winn et al. 2013), KAPTUR (Garrett et al. 2012), Orbital (Stainthorp, 2012) and RoaDMaP (2013) [see section 7.1. for more information about these JISC RDMI projects].

Researcher requirements:       
      a) For active data
·         Direct capture of data (and metadata) from instrument.
·         As much automated metadata annotation as possible, such as project level metadata (researcher identity and grant information) imported from the CRIS.
·         Network that provides adequate storage (personal and project) which is regularly backed up with speedy access to large data volumes.
·         Secure, authenticated access mechanisms are required, especially for sharing sensitive data; usually involves institutional authentication mechanisms (Shibboleth [4.9.6]).
·         Ability to share data with collaborators inside and outside the institution (‘Academic Dropbox’).
·         Mechanisms for secure data destruction.
·         Mechanisms for data transformation as required for data curation (such as anonymisation, aggregation and format transformation).

b)      Depositing archive data
·         User friendly data upload facility (like Dropbox [4.4.6]).
·         Customisable workflows for creating or importing metadata and uploading file.
·         Simple process for ingest of large data collections (multiple files) and association of collections with single metadata record (dataset record).
·         Controlled lists for some metadata fields.
·         Support for versioning of datasets.
·         Clear choice of license options.
·         Specify granular access rights to files at data object and collection level.
·         Embargo options for metadata and files.
·         Mechanisms for secure data destruction.

      c) For data discovery and reuse
·         Effective search and discovery mechanisms, using subject-specific terminology. Controlled vocabularies of keywords with auto-complete function.
·         Enable immediate access to datasets.
·         Access to datasets held outside the repository.
·         Support access to very large datasets.
·         Means of access to restricted data, where the metadata is visible; a ‘contact owner’ button.
·         Linking dataset to context / reuse metadata or data documentation – describing the process of data generation.
·         Related data and research publications indicated and linked to.
·         Support for granular access to data and associated metadata.
·         Visualisation and data analysis tools to give summary data or overview of data. Support query and processing of data on the repository server rather than after download.
·         Support for free tagging - adding discipline specific tags or metadata to datasets.
·         Federated catalogues allowing searching across multiple institutions.
·         Advice on data citation.
·         Citation data produced, demonstrating impact.

Additional RDM service requirements:
·         Customisable metadata schema.
·         Support multiple ingest protocols.
·         Staged deposit workflow – allows administrative area for quality check / validation.
·         Enable selective metadata harvesting.
·         Enable extraction of metadata and data in open format.
·         Support open standards and exposure of metadata.
·         Support multiple content licensing – exposed clearly.
·         Support technical metadata.
·         Support generation of persistent unique identifiers.
·         Support open methods of authentication.
·         Ability to remove data to access controlled area, a dark archive for embargoed data
·         Ability to delete data, generating a tombstone reference.
·         Access to metadata through library catalogue – OAI-PMH [4.8.3] endpoint required.
·         Support reporting – analysis of repository content, download and view metrics.
·         Enable creation and retrieval of an audit trail, reporting management actions.


2.3.   Institutional considerations

Expediency may perhaps determine the development of the Institutional RDM technical infrastructure. In the current climate of budget constraint and with the need to demonstrate value for money, there should be a focus on appraising the systems currently in place, and determining whether these may be modified to fit the proposed infrastructure. Modification of existing components will require local expertise or employment of developers, often the more expensive aspect of system implementation. Thus, work will be needed in costing the options available: building upon and integrating existing components, or otherwise implementing a new fully-integrated system, possibly a proprietary system, replacing existing components where necessary.

The least the institution needs to do for the development of the RDM technical infrastructure:
  • Implement institutional policy – Additional infrastructure and services for research data management, to be developed in consultation with researchers.’ Therefore a research data audit is recommended to determine researcher practices.
  • Fulfil Funder requirements – ‘Research organisations will ensure that appropriately structured metadata describing the research data they hold is published...’ Therefore a data catalogue is required by 1st May 2015.
  • Promote and facilitate good RDM practice. Training and guidance resources need to be developed.
  • Select sustainable, inexpensive, open options (open for interoperability and sustainability). Business cases will need developing for the various options available.
  • Take into account projected future requirements. This involves a consideration of risks to services through the removal of funding (for example the AHDS[iv] data centre no longer received funding after 2008, so stopped functioning).


2.4.   Requirements gathering methods

In developing the institutional RDM strategy, the DCC recommends using both requirements-gathering and gap analysis methods (Jones et al. 2013). The DCC provide a number of tools for the purpose and have published a case study detailing the use of these tools (Rans and Jones, 2013). 

A number of UK institutions have used the Data Audit Framework (DAF)[v] developed by JISC and HATII for requirements gathering. The DAF provides a set of survey methods, questionnaire and interview frameworks in order to identify, locate and describe research data assets and determine how they are being managed. The AIDA Toolkit[vi] has also been developed for institutional self-assessment of the readiness and capabilities for management of digital assets and digital preservation.

The Collaborative Assessment of Research Data Infrastructure and Objectives (CARDIO)[vii] is a benchmarking tool for RDM strategy development developed from key aspects of DAF and AIDA and other tools. The DCC recommend using CARDIO in conjunction with the other tools, the emphasis being on strategic planning and identifying gaps between the current situation and best practice. 





[ii] Datacite metadata schema 3.0 http://schema.datacite.org/meta/kernel-3/ [See section 4.9.1.]
[iv] Arts and Humanities Data Service (AHDS) http://www.ahds.ac.uk/
[v] Data Audit Framework (DAF) http://www.data-audit.eu/index.html
[vii] Collaborative Assessment of Research Data Infrastructure and Objectives (CARDIO) http://cardio.dcc.ac.uk/

RDM Technical Infrastructure Review - Introduction

1. Introduction

The purpose of this report is to indicate options available for the development of a technical infrastructure to support research data management (RDM) at the University of Sheffield. RDM, its situation within academic research and recent drivers towards change are defined. The processes involved in RDM and the elements of the supporting technical infrastructure are examined. The local context, of RDM technical infrastructure at the University of Sheffield and collaborating institutions, is explored.  The range of technical infrastructure components available and evaluations of these components are reviewed. Instances of fully-functioning RDM technical infrastructure and many of the recent research projects that developed and piloted RDM technical infrastructure components are briefly described. Finally, recommendations for suitable technical infrastructure components are proposed.

1.1.   Research Data Management

The research data collected to test a research assertion must be managed in an appropriate manner to be considered good research practice. Good RDM practice is now required of researchers by many research institutions and by most research funders. Increasingly research funders are demanding long-term curation of some of the data resulting from the research they fund, so that they may be available for re-use. The value of those data and the impact of the original research are increased by re-use. RDM may be considered to involve three broad areas of activity:-
  • Data management planning, during the research proposal and grant application stage.
  • Looking after ‘live’ or ‘active’ data as they are collected, processed, shared and stored.
  • Data Stewardship - Long-term curation of research data and data publishing, making data discoverable and reusable.

1.2.   Research data management drivers 
JISC[i] have supported the development of RDM practice over the last fourteen years by funding projects involving HEIs through a number of programmes, in particular the Managing Research Data Programmes 2009-11[ii] and 2011-13[iii]. The DCC[iv] was established in 2004 with JISC funding, to support expertise and practice in RDM. Since 2011 the DCC have offered tailored support in the development of policy, services and infrastructure for Higher Education Institutions, and are the foremost source for information and advice in the development of RDM infrastructure (Jones et al. 2013). Infrastructure refers to the hardware, software and human resources necessary to support the RDM services and processes. This report focuses on the technical infrastructure, the software and hardware components available.


The need for good RDM practice is recognised by all stakeholders involved in the research process. These include: 
  • Researchers, who may be part of a research project team, which may include members of many different institutions. Researchers need to secure their data against loss or unauthorised access. Making data available to reuse allows verification, promotes integrity and increases research impact.
  • The research institution (may be a HEI) or body employing the researcher and providing the facilities. Research data may be considered part of an institution’s special collections. Institutions will also wish to minimise risk to the data and damage to their reputation.
  • The research funder (usually a research council, charity or a HEI) who may mandate RDM procedures such as the creation of a Data Management Plan (DMP) and the deposit of data to repositories. Research Funders may support facilities for data curation such as data centres.  Funders wish to increase the return on their funding, through the reuse of data.
  • Governments, who fund research councils and other funding bodies, are concerned to derive as much value as possible from publicly funded research.
  • Publishers of research papers, who may publish the underlying data, seeking to add value to the publication process. 
The EPSRC policy framework on research data[v], published in May 2011, puts forward the EPSRC expectations[vi] of organisations receiving EPSRC funding, concerning the management and provision of access to EPSRC funded research data. These nine expectations were developed from seven guiding principles[vii] which are aligned with the RCUK common principles on data policy[viii]. Institutitions in receipt of EPSRC funding are expected to be fully compliant with these expectations by 1st May 2015. In terms of RDM Infrastructure, the pertinent expectations (EPSRC, 2013) are:

“Research organisations will ensure that appropriately structured metadata describing the research data they hold is published (normally within 12 months of the data being generated) and made freely accessible on the internet; in each case the metadata must be sufficient to allow others to understand what research data exists, why, when and how it was generated, and how to access it”

“Research organisations will ensure that EPSRC-funded research data is securely preserved for a minimum of 10-years from the date that any researcher ‘privileged access’ period expires or, if others have accessed the data, from last date on which access to the data was requested by a third party;”

“Research organisations will ensure that effective data curation is provided throughout the full data lifecycle... The full range of responsibilities associated with data curation over the data lifecycle will be clearly allocated within the research organisation, and where research data is subject to restricted access the research organisation will implement and manage appropriate security controls;”
The University of Sheffield Research Data Management Policy[ix] was developed in response to the RCUK principles and EPSRC expectations. Of the eight points of policy, the following (The University of Sheffield, Research and Innovation Services, 2014) are particularly relevant for RDM Infrastructure:

The primary responsibility for effective research data management during the course of research projects lies with lead researchers. However, all researchers, including postgraduate and undergraduate students undertaking research, have a personal responsibility to manage effectively the data they create.”

“Unless the terms of research grants or contracts provide otherwise, data generated by research projects are the property of the University of Sheffield. Researchers should exercise care in assigning rights in data to publishers or other external agencies.”

The University will provide support for research data management, including… ...Additional infrastructure and services for research data management, to be developed in consultation with researchers.”

The research institution is the body responsible for providing the researcher with facilities for research and is therefore responsible for providing the researcher with the necessary infrastructure and services to support RDM. Design of this infrastructure must be based upon the researcher workflow, so as not to burden the researcher with additional work or by changing their practices, and where possible, making RDM processes virtually automatic and invisible to the researcher.

1.3.   Research data lifecycle

Data collected during a research project will include the research data themselves; experimental, observational, modelled data etc. together with the metadata that describes these data in detail and documentation describing the context of the research, details of the research project and the processes involved. Data here will be defined as the numerical and textual information collected by analysis or measurement from the research samples or objects – but not the samples or objects themselves. For example, a collection of images or of tissue samples will not be considered data until there is textual or numerical information, such as identifiers (names or ID numbers), descriptions and relationships, associated with them. In this document, data refers to digital data, although the same management principles apply to analogue data formats. However, it is best practice to digitise research data, making its discovery and reuse easier.

From the initial drawing-up of a research proposal and grant application, data will be collected and managed. This includes documentation about the project, people and bodies involved, grant application, data management plans, experimental protocols, possibly test data or data collected from previous projects for re-use and a literature review or bibliography.

Once underway, a research project will collect or create raw data, which, during the project, will usually be processed to create derived or processed data. There may be many different iterations of processing, resulting in many sets of derived data. Eventually a set of ‘results’ data will be selected as the basis of the research publication(s) output by the project. All these sets of data can be considered active data, which will need to be quickly accessible and easily shared between collaborators. All these sets of active data will need appropriate documentation to describe the processes involved in their creation and modification.

After the project has finished, researchers will need to select data for curation on a long-term basis. This may have been decided in agreement with the research funder during the initial planning stage of the project. Curation in the context of RDM, refers to archiving, preservation and adding value through transformation and reuse. These archive data selected for curation, may need further processing (validation, cleaning, anonymisation or redaction) before submitting to an appropriate repository. The associated metadata will be needed to provide the necessary information for citation and re-use. There may be the facility to add new metadata or documentation, generated by data reuse, to the curated dataset. Data not required for curation needs to be disposed of in an appropriate manner.

1.4.   Data documentation, metadata and data collections

Data need to be documented to be understood and managed. Data documentation indicates the conditions and processes involved in the creation or collection of the data, the processing of the data and the context of the research. Detailed documentation is essential for verification and reuse. Adequate data documentation is necessary to determine provenance, licensing and access arrangements and preservation requirements. Research data need to be documented at three levels:
  • Project level – providing an overview of the research context and design.
  • File level – describe the relationships between files or database tables.
  • Item level – describing, for example, the meaning of a variable in a table.                                             (Research Data Mantra, 2014UKDA, 2014)
Metadata are a highly structured subset of core data documentation. Metadata are structured so that they may be indexed and stored within a database, thereby facilitating data organisation and discovery, and machine to machine interoperability. By considering its function, metadata may be divided into three layers:
  • Core metadata (Datacite, 2011, p. 8) or Catalogue metadata - creator name, publisher, title and an identifier are required for discovery and correct citation of the data. This could possibly include some subject description or classification details.
  • Detail metadata or Administrative metadata – provides generic dataset description. This includes access, preservation and technical metadata, and is required for the long-term curation of the data. This will include more detailed classification / subject description.
  • Discipline specific metadata (also known as Reuse metadata) – This documents aspects of the dataset that will be of interest to researchers wishing to validate the research process or re-use the data. This will consist of experimental protocols, instrument settings, and relationships with other elements of the dataset, other files within a data collection or other data collections. This will provide very detailed classification / subject description, providing the fine-grained attributes of data necessary for accurate discovery and location of elements within a dataset. Discipline specific information is frequently held in unstructured formats, so could be considered data documentation rather than metadata.                                                                                                                             (Ensom, 2013 and IDMB, 2011)

Data collections are typically organised by reference to a particular survey or research topic and may cover a specific geographic area and time period. The UKDA defines a data collection as typically comprised of three components: data, documentation and metadata. Code is occasionally considered a fourth component (Ensom and Corti, 2012, p. 3).

1.5.   Data repository or Data registry?
A repository is a content management system which may be considered to consist of three elements – a user interface front-end, and a database layer and storage layer back-end. The database holds records of entities, consisting of metadata elements as a series of fields. The storage layer, containing the actual data bitstreams, may be a file system on the repository server, or a file system on a local server or a remote server that is independent of the repository system. This may include cloud storage or a hybrid storage system. Usually in a repository system, only metadata records are held within the database, not the data objects themselves. This is due to the larger size of data objects which results in slower indexing / access speeds. Storage designation is handled by a storage controller or storage resource broker.

Different repository systems may be configured for different organisational structures. Repositories may range in the granularity of data described: The entity described by a repository record, the data object, may be a single row within a database or spreadsheet, a single file, or a collection of interrelated files constituting a dataset. In the Essex ePrints [6.1.4][x] context, the ‘eprint’, the key entity, is the ‘data collection’ which consists of a set of metadata and files (Ensom and Wolton, 2012). Datasets may be grouped as collections or groups, the ‘User’ entities may be grouped as a ‘Community’.

A ‘Data registry’, or ‘Data catalogue’ or ‘Metadata store’, is a repository system that holds only metadata records. The data themselves are held on a local or remote file storage system, so the metadata record points to the data store filepath or URL. By using the appropriate metadata standards, metadata records may be exchanged between registries, giving rise to the possibility of national and international registries. Repository systems were originally designed to curate textual digital objects, but are now being modified to curate digital objects of all formats and sizes, with the view of extending their purpose to curate research data (Gutteridge, 2010). As they have been designed to manage, curate and publish research outputs, they may perhaps provide the ideal platform for a catalogue for the data underlying the research outputs.

The ANDS[xi] strategy has been to separate the data storage function from the cataloguing function, the dissemination function and access control are provided by the metadata store (ANDS, 2011b).

1.6.   Research data ecology
In considering the implementation of an Institutional RDM technical infrastructure, it can be useful to consider the ecological approach (Robertson et al. 2008), to gain a better understanding of the interactions between repositories and services. The local infrastructure does not exist in a vacuum and must interact with, and is dependent on, a diverse range of entities and processes in the information ecosystem.

In determining the place of a repository, registry or other infrastructure component in the overall information ecosystem, it may be helpful to identify a range of components by considering their data storage coverage and specialisation (ANDS, 2011a), which may be:
·         Local – personal, project or departmental server for active data storage.
·         Storage associated with an instrument or facility.
·         Institutional storage for active data – networked filestores.
·         Institutional repository for archive data.
·         Multi-institutional project data storage (CARMEN [6.2.5] for example).
·         Research council data centre – for archive data including longitudinal data.
·         Discipline based repository – for active and archive data. National or International coverage.

Metadata storage will have a similar range (ANDS, 2011b), which may be:
  • Local, project and instrument based metadata store – spreadsheets or databases associated with the data.
  • Institutional data registry or repository.
  • National data registry (ANDS).
  • Discipline based metadata store – may be international, national or multi-institution based.

A number of practitioners have suggested that, regarding research data curation, “the Institutional repository is the repository of last resort” (Haywood, 2013), since discipline based repositories are better configured for the types of data and specialised metadata formats associated with the research community they serve. However the importance of the institutional research data services (tier 3) in the hierarchy of rising value and permanence (Figure 1. below)  is emphasised in the Royal Society report ‘Science as an open enterprise’ (2012).



Figure 1. The data pyramid - a hierarchy of rising value and permanence (Royal Society, 2012)

This is reiterated by Simon Hodson (2012), who maintains that Institutional research data services are essential because:
  • The institution is where the data are created and can be captured. Institutions implementing RDM infrastructure will make data discovery and curation possible.
  • Joining the gulf between curated data in national / international data services and uncurated and inaccessible data in individual or project collections.
  • Elevating data to national / international data services from temporary and inaccessible individual collections. Important data collections may emerge as they become discoverable.

The data catalogue component of an institutional infrastructure would ideally adhere to the formats and schemas used by a national data registry under development. The existence of tools for interoperability and for deposit to the major data centres should also be a consideration in the selection of a repository system.

1.7.   Development of RDM services

The DCC have created a guide to developing RDM services at HEIs (Jones et al. 2013), which breaks down the development, process into a number of components, as visualised in Figure 2. 


Figure 2. The components of an RDM service as envisaged by the DCC (Jones et al. 2013, p. 5)

The approach to the development of RDM Services for HEIs, recommended by the DCC involves:
  • Assembling a steering group composed of senior representatives of stakeholder group – senior researchers and institutional support service managers.
  • Appointing an RDM service development group to undertake the work.
  • Carrying out a gap analysis, to determine gaps between current position and the aimed-at future position, and requirements gathering surveys to determine stakeholders needs.
  • Development of RDM policy and strategy. Developing a policy first may be useful as a motivating factor, but may lead to problems if the proposed infrastructure and services cannot be realised. Alternatively, the development of policy may be subsequent to the defining of strategy. 
  • Designing services to meet local and external needs – putting in place the infrastructure required to support these services.
  • Piloting these services to test that they are fit for purpose. 
The DCC have published a case study detailing the development and implementation of an RDM strategy (Rans & Jones, 2013). Design of the technical infrastructure required to support the RDM Service is considered below.




[i] Joint Information Services Council (JISC) http://www.jisc.ac.uk/
[iv] Digital Curation Centre (DCC) http://www.dcc.ac.uk/
[v] Engineering and Physical Sciences Research Council (EPSRC) policy framework on research data http://www.epsrc.ac.uk/about/standards/researchdata/Pages/policyframework.aspx
[viii] Research Councils UK (RCUK) common principles on data policy http://www.rcuk.ac.uk/research/datapolicy/
[ix]The University of Sheffield Research Data Management Policy  http://www.shef.ac.uk/ris/other/gov-ethics/grippolicy/practices/all/rdmpolicy
[x] Essex Research Data http://researchdata.essex.ac.uk/ [see section 6.1.4 for information]
[xi] Australian National Data Service (ANDS) http://ands.org.au

RDM Technical Infrastructure Review - Executive Summary & Contents

A Review of Options for the Development of Research Data Management Technical Infrastructure at the University of Sheffield

A Report to the University of Sheffield Research Data Management Service Delivery Group

John Lewis  29/04/14


Executive Summary

This report reviews the options available for the development of a technical infrastructure, the software and hardware systems, to support Research Data Management (RDM) at the University of Sheffield. The appropriate management of research data throughout the data lifecycle, during and after the research project, is considered good research practice. This involves data management planning during the research proposal stage; looking after active data, its creation, processing, storage and access during the project; and data stewardship, long-term curation, publishing and reuse of archive data after the end of the project.

Good RDM practice benefits all stakeholders in the research process: Researchers, will secure their data against loss or unauthorized access, and may increase research impact through publishing data; Research institutions may consider research data as ‘special collections’ and will need to minimise risk to data and damage to reputation; Research Funders wish to maximise the impact of the research they fund by enabling reuse; Publishers may wish to add value to research papers by publishing the underlying data.

Many research funders now mandate RDM procedures, particularly Data Management Planning (DMP), and the UK research councils policies have contributed to the RCUK common principles on data policy. Notice must be taken of the EPSRC Expectations of organisations receiving EPSRC funding. These include the requirements that the organisation will:
  • Publish appropriately structured metadata describing the research data they hold - therefore the institution must create a public data catalogue.
  • Ensure that EPSRC-funded data is securely preserved for a minimum of ten years – therefore the institution must create a data archive.
  • Ensure that effective data curation is provided throughout the full data lifecycle – therefore the institution must provide the necessary human and technical infrastructure required. 
Institutions in receipt of EPSRC funding are expected to be compliant with these expectations by 1st May 2015. The University of Sheffield Research Data Management Policy was developed in response to the RCUK principles and EPSRC expectations. This states that the University will develop infrastructure and services to support research data management in consultation with researchers.

The local infrastructure does not exist in a vacuum and interacts with, and is dependent upon a range of other services and processes in an information ecosystem. At one end of the continuum of research data curation is the local storage of data and metadata (data identification, description and documentation), usually accessible to the project team only. At the other end are international discipline-based data repositories or national data centres that publish research data, facilitating its discovery and access. Research institutions lie in the middle of this continuum and provide the means to move research data and metadata from their local, unpublished state to an international published state. Some institutions now publish research data, either by modifying the institutional repository (IR) to accommodate datasets in addition to research papers or through a data repository, being a new instance of repository system running alongside the IR. However, discipline-based repositories are considered the most appropriate facility for data publishing, due to their configuration for the data types and metadata formats associated with the research community they serve.

The repository is here defined as a software system composed of three layers – a user interface, a database holding metadata records, and a storage layer holding the actual research data bitstreams. In some implementations, often known as data registries, data catalogues or metadata stores, the repository holds only the metadata records and links to the data stored elsewhere.

This report focuses on the outcomes of projects at UK HEIs funded by the JISC ‘Managing Research Data’ programmes 2009-11 and 2011-2013. Generally the infrastructure architectures examined have been developed in response to the functional requirements derived from researcher workflows. The major functional components of the RDM technical infrastructure for the institution are:
  • Metadata capture system – in order to identify, describe and document the research data as they are created, captured and processed and record the context, conditions, variables and instrument settings. This may be accomplished manually, by the researcher filling in forms, or automatically, concurrent with data capture, by using appropriate equipment.
  • Active research data management system – Active research data needs to be accessed rapidly, may require large computational resources and may require stringent security and access arrangements. A number of collaborative systems and virtual research environments have been developed to fulfil these requirements. These can be considered to comprise of a filestore and a data registry (sometimes known as a metadata store or asset registry).
  • Research Data Repository – will be an appropriate place for preservation and publishing of archive research data for which there is no discipline-based repository or data centre available. The catalogue and archive functions of the repository may be separated.
  • Research Data Catalogue – holds the metadata records of published (but not necessarily open access) research data. The data themselves may be held in a discipline-based data repository outside the institution or in an institutional data archive.
  • Research Data Archive – preserves data not, or not yet, submitted to discipline-based data repositories. The associated metadata records will be held in the research data catalogue.
  • Current Research Information System (CRIS) – manages the metadata associated with researcher identity, project information, research costing, grant applications and awards.
These components may overlap in function, but need to be interoperable to provide seamless RDM. Alternative approaches to an infrastructure composed of diverse components, where ensuring interoperability may be problematic, are provided by data grids and micro-services. The technical infrastructure must fit into the researcher workflow, making RDM processes automatic and virtually invisible to the researcher as far as possible. This is so as not to burden the researcher with additional work or changes to their practice. Products, processes and practices that have been developed by a researcher community should be adopted, adapted and developed for the needs of other researchers, rather than new solutions developed. 

The choice of technical infrastructure components and the approach of implementation will need to be considered with regard to the infrastructure and expertise already present. Integrating and modifying existing components may be as expensive, in terms of development work, as implementing new infrastructure. Installing and configuring free open-source software may prove expensive in terms of development, compared with proprietary systems. At the University of Sheffield there is currently no system that adequately supports collaborative active data management – a virtual research environment, or ‘academic dropbox’. There is little information available regarding the use of data and metadata capture systems, although systems such as laboratory information management systems may be used by some research groups at the institution. It is feasible that the Symplectic CRIS may be configured for use as a research data registry. The ePrints institutional repository, WRRO, may be configured for use as a research data catalogue, but this relies on agreement between the WRUC member institutions. In order to implement an independent institutional research data repository, expertise will be needed to do the necessary development work.

A shared approach to RDM services is being investigated by the White Rose and N8 consortia (of which the institution is a member); a WR Research Data Catalogue and N8 shared data archiving service having been proposed. The development of RDM services delivered through the White Rose Grid and N8 HPC grid infrastructure need to be explored. Attention should be paid to the national research data service is being piloted through the DCC and JISC. The great benefits of the shared approach demand that support for collaborative projects establishing shared RDM services should be a priority.

This report briefly describes the eighty most commonly used components of RDM technical infrastructure at UK HEIs. The report describes evaluations, reviews and comparisons of these components, gives examples of established RDM services and highlights the recent projects at UK HEIs which were involved in developing these services. 

By way of conclusion, a number of recommendations are made regarding the choice of infrastructure components to be made and implementation strategy to be considered. These recommendations take into account the current situation of technical infrastructure at the institution and the constraints on time and cost. Attention is drawn to the development of shared RDM services with collaborating institutions. Finally the proposal is put forward that a number of technical infrastructure components of an integrated RDM service are first piloted with EPSRC funded research projects to ensure compliance with EPSRC expectations by 1st May 2015.


Contents

1.     Introduction (1)
1.1   Research data management (1)
1.2   Research data management drivers (1)
1.3   Research data lifecycle (3)
1.4   Data documentation, metadata and data collections (4)
1.5   Data repository or data registry? (5)
1.6   Research data ecology (5)
1.7   Development of RDM services (7)
2.     RDM Technical Infrastructure Architecture (9)
2.1   Technical infrastructure components (9)
2.2   Functional requirements (12)
2.3   Institutional considerations (13)
2.4   Requirements gathering methods (14)
3.     The University of Sheffield RDM Technical Infrastructure Considerations (15)
3.1   Local infrastructure components (15)
3.2   Consortia options (16)
3.3   Recent reviews of RDM service developments at Sheffield (19)
4.     Infrastructure Components (23)
4.1   Integrated systems and integrating components (23)
4.2   Repository platforms (23)
4.3   Archive data storage and digital preservation systems and services (26)
4.4   Active data management and collaboration platforms (27)
4.5   Catalogue software (29)
4.6   Current Research Information Systems (CRIS) and DMP tools (30)
4.7   Data capture and workflow management systems (30)
4.8   Data transfer protocols (33)
4.9   Identifier services and identity components (33)
4.10 Other software systems and platforms of interest (34)
5.     Reviews, Evaluations and Comparisons of Infrastructure Components (36)
6.     Active Institutional Infrastructure Examples (41)
6.1   UK institutional data repositories (41)
6.2   Discipline-based research data repositories hosted by UK HEIs (43)
6.3   Institutional and discipline-based research data repositories outside the UK (43)
7.     RDMI Project Outputs (46)
7.1   Outputs from the JISC RDMI 2011-2013 projects (46)
7.2   Outputs from the JISC RDMI 2009-2011 projects (49)
7.3   Outputs from other relevant projects (51)
8.     Conclusions and Recommendations (52)
9.     References (55)
9.1   Works cited in the text (55)
9.2   Index of entities noted in the text (61)

Wednesday, 19 February 2014

Datacite Metadata Generator

This is a single HTML form which can be used to generate DataCite Metadata Kernel 3.0 XML.
Metadata is generated by populating the text boxes and selecting values from drop-down lists. The results can be saved to a file. This tool has been created by Marcin Paluch and is available at https://github.com/mpaluch/datacite-metadata-generator.

The tool is described at  http://www.datacite.org/node/102