For the complete documentation index, see llms.txt. This page is also available as Markdown.

FAIR Benchmark - Institutional Repository Datasets

Introduction

FAIRsharing is collaborating with several organisations, including UKRN (via the ORP, see project page) and the University of Oxford (see our news item), to develop a FAIR benchmark for datasets held in institutional repositories. The benchmark has been created using the OSTrails Assess-IF framework. It uses a mixture of metrics specific to the needs of institutional repositories as well as reusing existing generic metrics, enabling communities to build upon established FAIR definitions wherever possible. This approach promotes consistency, comparability, and transparency, with all benchmark components, including metrics, tests, definitions, and justifications, being openly registered, discoverable, and reusable. Using this framework allows communities to explicitly define what FAIR means in their own context while remaining aligned with a broader ecosystem of shared FAIR assessment components.

The Institutional Repository Datasets Benchmark provides a structured framework for assessing and improving the FAIRness of metadata describing research datasets deposited in institutional repositories. It operationalises the FAIR Principles in a practical and transparent manner, supporting alignment with community-endorsed standards and good research data management practices.

The benchmark is intended for two primary audiences. Institutional repository teams can implement and run the associated assessments as part of repository workflows, while researchers depositing datasets can use the resulting feedback to understand the FAIRness of their records and identify opportunities for improvement. By supporting both repository managers and depositors, the benchmark promotes FAIR awareness and enables targeted FAIRification actions.

Designed for use with publicly accessible repository records, the benchmark supports consistent FAIR assessment across institutional repositories and research disciplines. It distinguishes between FAIR-enabling properties implemented by the repository itself and FAIR characteristics that depend on the deposited record, while clearly identifying where existing metrics have been reused, adapted, or newly developed to address the specific needs of institutional repositories.

Weighting and Compliance Measurement

This benchmark evaluates FAIRness through a collection of FAIR metrics aligned to the FAIR Principles. Each metric is implemented by one or more tests that return a pass, fail or indeterminate result. All metric results are retained and reported as part of the assessment record to ensure full transparency and traceability of the assessment process.

Metric-level scores are calculated first, to group together all tests linked to a single metric. Each possible test result (other than error) carries its own individual weight using a scale of 0 – 5, where 0 means that the test result does not influence the FAIR evaluation at all and 5 is used for the most important test results from the perspective of the community. Note that while negative weightings are allowed by the OSTrails framework, these have not been defined within this community. The scoring algorithm contains the weights to be applied for each pass, fail or indeterminate test result. Currently, passing results are the only ones that contribute positively its metric score, while failing, indeterminate or error results are weighted 0 and contribute no score. The relative contribution of individual metrics within a principle reflects the community priorities defined for this benchmark. The numerical scores and resulting compliance bands are not intended to function as a grade, ranking, or overall measure of dataset quality. The goal of the benchmark is to identify gaps in compliance by the repository and/or the submitter such that tailored guidance can be created.

Principle-level alignment is calculated from the combined weighted results of all metrics associated with that principle and are expressed as a percentage of the maximum possible score for that principle. These principle-level scores are then combined to produce category scores for the four FAIR dimensions (Findable, Accessible, Interoperable and Reusable). Each principle contributes according to the community-defined weighting of its constituent tests, that reflecting its relative importance within the context of institutional repositories. This allows communities to express priorities while maintaining a transparent and reproducible scoring methodology.

Finally, an overall FAIR result is calculated as the average of the four normalised F, A, I and R category scores. This ensures that each FAIR dimension contributes equally to the overall assessment, while allowing different communities to assign different relative importance to the principles that comprise each FAIR dimension.

Compliance Bands

For reporting and visualisation purposes, normalised scores are mapped onto four compliance bands: 25%, 50%, 75% and 100%. A score reaches a given band when the corresponding normalised score equals or exceeds that threshold. The 100% band is reserved for complete compliance and is only awarded when all metrics contributing to the relevant principle or FAIR category have passed.

These bands provide a simple and consistent representation of FAIR maturity across all principles and FAIR dimensions while preserving the underlying metric-level assessment results.

Repository-Managed and Submitter-Managed Metrics

During development of the benchmark, metrics were classified into two categories:

  • Repository-managed metrics, representing characteristics primarily determined by repository infrastructure, policies, or technical implementation.

  • Submitter-managed metrics, representing characteristics that can be directly influenced by the individual creating or editing a metadata record.

This distinction was introduced to support future development of more targeted FAIR assessment and improvement workflows. For example, repository-managed and submitter-managed scores could be reported separately, enabling users to distinguish between issues that require repository-level intervention and those that can be addressed through metadata curation.

The current version of the benchmark records this classification but does not use it during score calculation, aggregation, or visualisation. The repository-managed versus submitter-managed classification is therefore retained as benchmark metadata for future use. A current listing of which tests are considered repository-managed and submitter-managed metrics is available in the "banding model" tab of the benchmark's scoring algorithm.

Last updated

Was this helpful?