Tier 1 · UniversalTrust the number

Know the model’s error

Every machine-learning prediction comes with an error bar, stated or not. If your decision margin is smaller than that error, the prediction cannot decide it.

5 min read3 worked examplesStage 04 in the research flowFact-checked Oct 2026
Illustration: Know the model’s error
In short

Look up the model’s benchmark error for the exact property you are using before acting on a prediction.

Compare that error with your decision margin: the distance between the predicted value and your threshold.

Expect larger errors on chemistries unlike the training data, and for derived quantities such as stability.

What it is

Machine-learning models — graph neural networks, universal interatomic potentials, descriptor-based regressors — now predict formation energies, band gaps, elastic moduli and stabilities for materials that have never been calculated or made. Each model has been evaluated on a held-out test set, which gives an average error such as the mean absolute error (MAE), or a classification metric such as precision or hit rate for finding stable materials.

Knowing the model’s error means carrying that number alongside every prediction and asking a simple question: is the prediction far enough from my threshold that the model’s typical error could not flip the decision? If yes, the prediction is decision-grade for screening. If no, it is a hint that needs a higher-tier check.

Public benchmarks make this practical. Matbench provides standardised tasks for common properties with fixed train/test splits. Matbench Discovery evaluates how well models identify thermodynamically stable crystals in a prospective-style setting, reporting both regression errors and classification metrics.

Schematic diagram: Know the model’s error
At a glance: Know the model’s error. Schematic, not to scale.

Why it matters for R&D decisions

ML screening is cheap enough that teams can rank hundreds of thousands of candidates in hours. The risk is not that models are useless — many are very good — but that their predictions are read as exact. A ranked list looks equally confident at position 5 and position 500. Comparing each prediction with the model’s error is the step that separates ‘the model says so’ from ‘the model is accurate enough to say so here’.

The formula

decisive if |ŷ − threshold| > k × MAE;   otherwise escalate to a higher-tier check
ŷ
The model’s predicted value for the candidate
threshold
Your go/no-go value for the property (e.g. 25 meV/atom above hull)
MAE
The model’s mean absolute error on a benchmark task for the same property
k
A safety factor you choose; 1–2 is common. Larger k means fewer false positives but more expensive follow-up

MAE is an average; individual errors can be several times larger. If the model provides calibrated uncertainties, use those per-prediction instead.

How to apply it, step by step

  1. 1
    Find the right benchmark number

    Look up the model’s error on a benchmark for the same property and similar data. An MAE on formation energy does not tell you the error on energy above hull or on band gap.

  2. 2
    Define the decision margin

    For each candidate, compute how far the prediction is from your threshold. A candidate predicted at 5 meV/atom above hull with a 25 meV/atom threshold has a 20 meV/atom margin.

  3. 3
    Compare margin with error

    If the margin is comfortably larger than the model’s typical error, treat the prediction as decision-grade for screening. If not, escalate the candidate to DFT or experiment.

  4. 4
    Check whether the candidate is in-distribution

    Errors grow for compositions, structure types or elements under-represented in the training data. A model trained mostly on oxides is less reliable on intermetallics or nitrides.

  5. 5
    Track hit rate, not just MAE, for discovery tasks

    When the goal is to find stable materials, what matters is how many predicted-stable candidates are actually stable. A model can have a low MAE and still misclassify many materials near the hull.

Worked examples

Example 1

Stability screen with a universal interatomic potential

Illustration for the example: Stability screen with a universal interatomic potential

You use a universal ML potential to estimate energy above hull for 10,000 hypothetical compounds. Gate: ≤ 25 meV/atom. Suppose the model’s benchmark MAE on energy above hull is 30 meV/atom (hypothetical, but of the order reported for strong models).

  1. 01Candidate A: predicted 150 meV/atom. Margin = 150 − 25 = 125 meV/atom ≈ 4 × MAE. Reject with confidence.
  2. 02Candidate B: predicted 15 meV/atom. Margin = 25 − 15 = 10 meV/atom ≈ 0.3 × MAE. The model cannot decide.
  3. 03Candidate C: predicted −40 meV/atom (below the current hull). Margin = 65 meV/atom ≈ 2 × MAE. Promising, but a claim of a new stable phase needs DFT confirmation.
  4. 04Escalate B and C to DFT relaxation; skip A.
RESULTExpensive DFT is spent only on candidates where it can change the outcome.

The same model is decisive for some candidates and useless for others; the margin decides which.

Example 2

Why a small formation-energy error can still mislead on stability

Illustration for the example: Why a small formation-energy error can still mislead on stability

A model predicts formation energies with an MAE of about 30 meV/atom. You want to use it to decide which compounds are on the convex hull.

  1. 01Stability is the difference between a compound’s energy and the energy of competing phases. Most of the formation energy cancels in that difference.
  2. 02In DFT, systematic errors partly cancel because the compound and its competitors are computed the same way. ML errors are less correlated between chemically similar compounds, so they cancel less.
  3. 03Typical decomposition energies of real materials are only tens of meV/atom, comparable to the model’s error.
  4. 04Result: classification of stable vs unstable can be much worse than the formation-energy MAE suggests.
RESULTYou evaluate the model with a stability-specific benchmark (precision and recall for stable materials) instead of the formation-energy MAE.

Check the error of the quantity you decide on, not of the quantity the model was trained on.

Example 3

Band gap model on an out-of-distribution class

Illustration for the example: Band gap model on an out-of-distribution class

A band-gap model reports an MAE of 0.35 eV on its test set (hypothetical). You apply it to a family of hybrid organic–inorganic halides that make up a small fraction of the training data.

  1. 01Check training-data coverage: few examples of this structural family.
  2. 02Run the model on five halides with known experimental gaps; observed errors range from 0.2 to 0.9 eV (hypothetical).
  3. 03Mean absolute error on this family ≈ 0.5 eV, larger than the headline MAE.
  4. 04Use the family-specific error, not the headline number, for decisions on this class.
RESULTThe screen widens its acceptance window for this family and adds a measured-gap check before ranking.

Headline benchmark errors are averages over the benchmark’s data; your chemistry may be harder.

Which error metric for which decision

DecisionMetric to look upWhy
Rank candidates by a propertyMAE / RMSE on that propertyTells you how much rank order can shuffle
Find thermodynamically stable materialsPrecision, recall, F1 or discovery hit rate for stabilityStability is a classification near a threshold
Filter by a thresholdError compared with margin to thresholdOnly near-threshold candidates are at risk
Use on a new chemistryError on a held-out subset of similar chemistryErrors grow out of distribution

When to use it — and when not to

Use it when
  • Every time an ML prediction influences which candidates go forward.
  • When choosing between models for a screening campaign.
  • When presenting predicted properties to non-specialists.
  • When deciding how many candidates to verify with DFT or experiment.
Don’t rely on it when
  • As the only check for materials far outside the training data — measure or compute directly instead.
  • As a substitute for physical constraints: a prediction inside the error bar can still be physically impossible (e.g. a negative band gap).

Common mistakes

Quoting the headline MAE for a different property than the one used in the decision.
Find a benchmark for the exact quantity: energy above hull, not formation energy; experimental gap, not PBE gap.
Assuming benchmark error applies to novel chemistry.
Check how similar your candidates are to the training data; estimate error on a related held-out subset.
Treating the ranked list as equally confident everywhere.
Mark which candidates are within one or two MAE of the threshold and send them for verification.
Counting an ML model trained on DFT as independent confirmation of that DFT.
Use it to extend coverage, not to triangulate the same data it learned from.
Ignoring the training tier.
A model trained on PBE inherits PBE’s systematic errors, such as underestimated band gaps.

Applying it in Lattice Graph

LatticeGraph includes Matbench and Matbench Discovery benchmark data alongside the computed sources models are trained on, so you can check a model’s reported error and compare predictions with DFT values for the same materials.

  1. 01Before acting on a predicted value, look up the relevant benchmark task and note the model’s error for that property.
  2. 02Compare the predicted value with your threshold and mark near-threshold candidates as provisional.
  3. 03Where a DFT value exists in Materials Project, OQMD, AFLOW or JARVIS for the same material, compare it with the prediction.
  4. 04Record the prediction, the model, and its benchmark error in your evidence pack.
DATASETS
MatbenchMatbench DiscoveryMPtrjMaterials ProjectOQMDAFLOWJARVIS-DFT

Frequently asked questions

What is a ‘good’ MAE?

It depends entirely on your decision margin. An error of 30 meV/atom is excellent for rejecting compounds hundreds of meV above the hull and inadequate for deciding between two candidates 10 meV/atom apart.

Are per-prediction uncertainties better than benchmark MAE?

When they are calibrated, yes — they tell you which individual predictions are risky. Check the model’s calibration before relying on them.

Why does Matbench Discovery report classification metrics as well as errors?

Because discovering stable materials is a classification problem near a threshold. A model with modest energy errors can still have high or low hit rates depending on how its errors are distributed near the hull.

References & further reading

  1. [1]
    Dunn, A., Wang, Q., Ganose, A., Dopp, D., & Jain, A. (2020). Benchmarking materials property prediction methods: the Matbench test set and Automatminer reference algorithm. npj Computational Materials, 6, 138.
    Defines the Matbench tasks and evaluation protocol.
  2. [2]
    Riebesell, J., Goodall, R. E. A., Benner, P., Chiang, Y., Deng, B., Ceder, G., Asta, M., Lee, A. A., Jain, A., & Persson, K. A. (2025). A framework to evaluate machine learning crystal stability predictions. Nature Machine Intelligence, 7, 836–847.
    Matbench Discovery: prospective-style benchmark for ML stability prediction.
  3. [3]
    Bartel, C. J., Trewartha, A., Wang, Q., Dunn, A., Jain, A., & Ceder, G. (2020). A critical examination of compound stability predictions from machine-learned formation energies. npj Computational Materials, 6, 97.
    Shows that accurate formation energies do not guarantee accurate stability predictions.
  4. [4]
    Sun, W., Dacek, S. T., Ong, S. P., Hautier, G., Jain, A., Richards, W. D., Gamst, A. C., Persson, K. A., & Ceder, G. (2016). The thermodynamic scale of inorganic crystalline metastability. Science Advances, 2, e1600225.
    Shows that decomposition and metastability energies of real materials are typically tens of meV/atom.
Results are informational and should be validated by qualified professionals. See Terms of Service