Tier 1 · UniversalTrust the number

FAIR data

Data should be Findable, Accessible, Interoperable and Reusable — so every number in your work can be traced, cited, combined and legally reused.

5 min read3 worked examplesStage all in the research flowFact-checked Oct 2026
Illustration: FAIR data
In short

FAIR is a set of 15 guiding principles published in 2016 for making data usable by people and machines.

In practice: persistent identifiers, rich metadata, standard formats, clear licences and recorded provenance.

For R&D teams, FAIR is what makes an internal result auditable a year later and an external dataset safe to build on.

What it is

The FAIR Guiding Principles were published by Wilkinson and colleagues in Scientific Data in 2016. They describe four properties that make research data usable not only by the people who created it but by others, and by software: Findable (data and metadata have persistent identifiers and are indexed), Accessible (they can be retrieved by identifier using an open, standard protocol, and metadata remain available even if the data are not), Interoperable (they use shared vocabularies and formats), and Reusable (they carry a clear licence, detailed provenance and community standards).

FAIR does not mean open. Data can be FAIR and still restricted — what matters is that the conditions for access and reuse are explicit and machine-readable. A proprietary dataset with clear identifiers, metadata, and a stated licence is more FAIR than an open spreadsheet with no provenance.

Materials science has built substantial infrastructure around these ideas: repositories such as NOMAD and Materials Cloud archive calculations with metadata; the OPTIMADE specification gives many databases a common query API; and large computed databases publish identifiers, versions and licences for their entries.

A useful way to read FAIR is as a checklist for the questions a future user will ask. Where did this number come from? Can I get the original record? Can I combine it with my other data without guessing at units or conventions? Am I allowed to use it for what I want to do? If a dataset answers those questions without anyone having to email its author, it is doing most of what FAIR asks.

Schematic diagram: FAIR data
At a glance: FAIR data. Schematic, not to scale.

Why it matters for R&D decisions

Most of the cost of reusing data is spent working out what a number means: which structure, which method, which version, and whether you are allowed to use it. Without that information, results cannot be reproduced, merged datasets silently mix incompatible conventions, and commercial teams can end up building on data whose licence forbids commercial use. FAIR practices turn each value into something that can be traced, compared and defended in a review, a regulatory filing or a due-diligence process.

How to apply it, step by step

  1. 1
    Give every record a persistent identifier

    Keep the source database’s own ID (for example a Materials Project or OQMD entry ID, a COD ID, or a DOI) with each value. For internal samples, assign IDs that never change and never get reused.

  2. 2
    Record the metadata needed to interpret the value

    At minimum: the material and phase, the property and units, the method (functional or measurement technique), conditions (temperature, pressure), and the source version or retrieval date.

  3. 3
    Use standard formats and vocabularies

    Store structures in widely supported formats such as CIF, use consistent units, and prefer community schemas (for example OPTIMADE property names) when exchanging data between systems.

  4. 4
    Carry the licence with the data

    Record each dataset’s licence at row or dataset level and keep it through exports. Licences vary widely: some sources are CC BY, some are share-alike, and some large datasets are released for non-commercial use only.

  5. 5
    Record provenance from source to decision

    Document how each derived value was produced: which inputs, which transformation, which code version. A claim in a report should be traceable back to source records in one or two steps.

  6. 6
    Check FAIR-ness when you ingest, not when you publish

    The cheapest moment to capture identifiers, licence and method is when a dataset first enters your system. Reconstructing them later, after values have been merged and transformed, is slow and often impossible. Make ingestion fail loudly if a source has no licence or no stable identifiers.

Worked examples

Example 1

Anatomy of a reusable data row

Illustration for the example: Anatomy of a reusable data row

A team records a measured ionic conductivity for a solid electrolyte in its shared dataset. The values are illustrative.

  1. 01Value and units: σ = 1.2 mS/cm.
  2. 02Tier and method: measured, electrochemical impedance spectroscopy (EIS), 25 °C, pressed pellet.
  3. 03Source: dataset name plus the source’s own entry ID, and the DOI of the originating paper.
  4. 04Licence: the dataset’s licence (for example CC BY 4.0), recorded with the row.
  5. 05Version: dataset version and retrieval date.
RESULTAnyone can find the original, understand the conditions, check whether they can reuse it commercially, and cite it correctly.

A handful of extra fields turn a bare number into evidence.

Example 2

Catching a licence problem before it ships

Illustration for the example: Catching a licence problem before it ships

A product team wants to include predicted crystal structures from a large ML-generated dataset in a paid screening service.

  1. 01Check the dataset-level licence recorded at ingestion: it is a non-commercial licence.
  2. 02Trace which shortlist entries depend on that dataset alone: 40 of 300 candidates (hypothetical).
  3. 03Options: obtain a commercial licence, regenerate the structures independently, or exclude those entries from the paid product.
  4. 04Record the decision and keep the licence field on every exported row.
RESULTThe product launches with the restricted entries excluded, and the evidence pack shows why.

A licence that is not carried with the data will be lost at the first export.

Example 3

Reproducing a result a year later

Illustration for the example: Reproducing a result a year later

An internal report claimed a cathode candidate was 12 meV/atom above the hull (hypothetical). A year later, a reviewer queries the number.

  1. 01Without provenance: the database has since been updated with new correction schemes and new competing phases; the current value is 30 meV/atom and nobody knows which version the report used.
  2. 02With provenance: the report stored the database version, entry IDs, correction scheme and the list of competing phases.
  3. 03Recompute with the stored version to reproduce 12 meV/atom, then explain the change from new competing phases in the latest version.
RESULTThe original decision is shown to be correct for the data available at the time, and the update is understood.

Versioned provenance turns ‘the number changed’ from a credibility problem into an explained update.

FAIR in practice for materials data

PrincipleWhat it asksMaterials example
FindablePersistent IDs, rich metadata, indexedEntry IDs, DOIs, searchable by formula and property
AccessibleRetrievable by ID over a standard protocol; metadata persistREST APIs such as OPTIMADE; archived records
InteroperableShared formats and vocabulariesCIF structures, SI units, OPTIMADE property names
ReusableClear licence, detailed provenance, community standardsLicence per dataset, method and version per value

When to use it — and when not to

Use it when
  • When designing any internal dataset, lab notebook schema or data pipeline.
  • When ingesting external datasets for analysis or commercial products.
  • When preparing evidence for a decision, publication, patent filing or due diligence.
  • When training ML models whose provenance you may later need to explain.
Don’t rely on it when
  • As a reason to delay all work until metadata are perfect — start with IDs, units, method, source and licence, and improve over time.
  • As a synonym for open data: FAIR can apply to confidential data with controlled access.

Common mistakes

Dropping the licence or version on export.
Make licence and version mandatory columns in every export template.
Replacing source IDs with internal row numbers.
Keep the original identifiers alongside any internal IDs so records can be traced back.
Storing values without method and conditions.
Record the functional or measurement technique and the temperature and sample form for every value.
Assuming ‘publicly downloadable’ means ‘free for commercial use’.
Read and record each licence. Non-commercial and share-alike terms have real consequences for products.
Merging datasets that use different conventions into one column.
Keep a source and method column for every merged value, and normalise conventions (units, sign, reference states) explicitly in a documented step.

Applying it in Lattice Graph

LatticeGraph keeps row-level provenance — source, source identifier and, where available, method — for values drawn from its sources, so that provenance can travel into your evidence pack and decision documents.

  1. 01When you shortlist a material, check each value’s provenance and source identifier.
  2. 02Note the licence of each source before using its data in a commercial deliverable.
  3. 03Export with provenance included, rather than copying bare values into a spreadsheet.
  4. 04Keep the export date so the result can be reproduced against the same data later.
DATASETS
Materials ProjectOQMDAFLOWJARVIS-DFTCODNOMADMaterials Cloud

Frequently asked questions

Does FAIR require making our data public?

No. FAIR is about clear identifiers, metadata, access conditions and licences. Confidential data can be FAIR within an organisation.

What is the minimum useful metadata for a materials value?

Material and phase, property and units, method, conditions, source identifier, licence, and version or retrieval date.

How does OPTIMADE relate to FAIR?

OPTIMADE is a common API specification that lets the same query run across many materials databases. It supports the Accessible and Interoperable principles.

Is FAIR only for computational data?

No. It applies equally to experimental data. For measurements, the most commonly missing metadata are sample preparation, phase purity, measurement conditions and instrument settings — exactly the fields needed to compare results between labs.

References & further reading

  1. [1]
    Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018.
    The original statement of the FAIR principles.
  2. [2]
    Andersen, C. W., Armiento, R., Blokhin, E., et al. (2021). OPTIMADE, an API for exchanging materials data. Scientific Data, 8, 217.
    Common API specification implemented by many materials databases.
  3. [3]
    Draxl, C., & Scheffler, M. (2018). NOMAD: The FAIR concept for big data-driven materials science. MRS Bulletin, 43, 676–682.
    How FAIR principles are applied to computational materials data.
  4. [4]
    Talirz, L., Kumbhar, S., Passaro, E., et al. (2020). Materials Cloud, a platform for open computational science. Scientific Data, 7, 299.
    Repository and provenance infrastructure for computational materials science.
Results are informational and should be validated by qualified professionals. See Terms of Service