For the complete documentation index, see llms.txt. This page is also available as Markdown.

FAIR Benchmark - Institutional Repository Datasets

Introduction

FAIRsharing is collaborating with several organisations, including UKRN (via the ORP, see project page) and the University of Oxford (see our news item), to develop a FAIR benchmark for datasets held in institutional repositories. The benchmark has been created using the OSTrails Assess-IF framework. It uses a mixture of metrics specific to the needs of institutional repositories as well as reusing existing generic metrics, enabling communities to build upon established FAIR definitions wherever possible. This approach promotes consistency, comparability, and transparency, with all benchmark components, including metrics, tests, definitions, and justifications, being openly registered, discoverable, and reusable. Using this framework allows communities to explicitly define what FAIR means in their own context while remaining aligned with a broader ecosystem of shared FAIR assessment components.

The Institutional Repository Datasets Benchmark provides a structured framework for assessing and improving the FAIRness of metadata describing research datasets deposited in institutional repositories. It operationalises the FAIR Principles in a practical and transparent manner, supporting alignment with community-endorsed standards and good research data management practices.

The benchmark is intended for two primary audiences. Institutional repository teams can implement and run the associated assessments as part of repository workflows, while researchers depositing datasets can use the resulting feedback to understand the FAIRness of their records and identify opportunities for improvement. By supporting both repository managers and depositors, the benchmark promotes FAIR awareness and enables targeted FAIRification actions.

Designed for use with publicly accessible repository records, the benchmark supports consistent FAIR assessment across institutional repositories and research disciplines. It distinguishes between FAIR-enabling properties implemented by the repository itself and FAIR characteristics that depend on the deposited record, while clearly identifying where existing metrics have been reused, adapted, or newly developed to address the specific needs of institutional repositories.

Weighting and Compliance Measurement

This benchmark evaluates FAIRness through a collection of FAIR metrics aligned to the FAIR principles. Each metric is implemented by one or more tests that return a pass, fail or indeterminate result. All metric results are retained and reported as part of the assessment record to ensure full transparency and traceability of the assessment process.

Each metric contributes equally to the principle-level assessment for the FAIR principle to which it belongs. A passing metric contributes the maximum available score for that metric, while a failing metric contributes no score. Because different FAIR principles are represented by different numbers of metrics, raw scores are normalised to a common 0-100 scale. This ensures that principles containing many metrics do not contribute disproportionately to higher-level FAIR scores. The numerical scores and resulting compliance bands (see below) are not intended to function as a grade, ranking, or overall measure of dataset quality. The goal of the benchmark is to identify gaps in compliance by the repository and/or the submitter such that tailored guidance can be created.

Principle-level scores are calculated from the combined results of all metrics associated with that principle. Each principle score is normalised to a percentage of the maximum possible. These principle-level scores are then aggregated to produce category scores for the four F, A, I and R dimensions. To avoid bias towards principles containing larger numbers of metrics, each FAIR dimension is calculated as the average of its constituent normalised principle scores rather than from raw metric totals.

An overall FAIR score is calculated as the average of the normalised F, A, I and R category scores. This approach ensures that each FAIR dimension contributes equally to the overall assessment, regardless of the number of metrics associated with that dimension.

Compliance Bands

For reporting and visualisation purposes, normalised scores are mapped onto four compliance bands: 25%, 50%, 75% and 100%. A score reaches a given band when the corresponding normalised score equals or exceeds that threshold. The 100% band is reserved for complete compliance and is only awarded when all metrics contributing to the relevant principle or category have passed.

These bands provide a simple and consistent representation of FAIR maturity across all principles and dimensions while preserving the underlying metric-level assessment results.

Repository-Managed and Submitter-Managed Metrics

During development of the benchmark, metrics were classified into two categories:

  • Repository-managed metrics, representing characteristics primarily determined by repository infrastructure, policies, or technical implementation.

  • Submitter-managed metrics, representing characteristics that can be directly influenced by the individual creating or editing a metadata record.

This distinction was introduced to support future development of more targeted FAIR assessment and improvement workflows. For example, repository-managed and submitter-managed scores could be reported separately, enabling users to distinguish between issues that require repository-level intervention and those that can be addressed through metadata curation.

The current version of the benchmark records this classification but does not use it during score calculation, aggregation, or visualisation. All metrics contribute equally to the principle-level FAIR scores described above. A current listing of which tests are considered repository- and which submitter-managed metrics is available in the "banding model" tab of the benchmark's scoring algorithm.

Last updated

Was this helpful?