# What It Means to Certify a Machine-Learning Model

Canonical: https://yourcloudblog.com/blog/what-it-means-to-certify-a-machine-learning-model/
Published: 2026-09-30
Author: YourCloudGroup · Publisher: YourCloudGroup · Category: Model Certification
Tags: model certification, machine learning, model governance, data leakage

> An accuracy figure is a claim. Certification is the evidence behind it: how we check that a model earned its score honestly, and why we would rather say 'not assessed' than guess.

**Short answer:** A certified model must never gain accuracy from information that would not exist when the prediction is actually requested. In YourCloudGroup Model Manager a certificate is fourteen independent, evidence-based claims — each CERTIFIED, NOT ESTABLISHED, NOT ASSESSED, NOT APPLICABLE or FAILED — not one pass/fail stamp.
Every model comes with a number. 94% accuracy. AUC of 0.98. A lift of 16× over random.
The number is where the conversation usually ends, and that is the problem. A score is a
**claim**, and most claims in machine learning are never checked.

The failures are rarely dramatic. A column that is only filled in *after* the outcome is
known slips into training, and the model learns to read the answer. The same test rows
get scored again and again until the model has, in effect, been tuned on them. A model is
retrained and quietly comes out different. None of these show up in the headline figure.
All of them make it worthless.

Certification in [YourCloudGroup Model Manager](https://yourcloudblog.com/model-manager/) exists to answer one question before anyone relies
on a number:

> **A certified model must never gain accuracy from information that would not exist
> when the prediction is actually requested.**

Everything below serves that rule.

---

## Certification starts before the model exists

Most tools evaluate a model after it has been built. By then it is too late to ask the
important questions, because the data the model learned from has already shaped it.

We put the integrity checks where the experiment is defined: at the moment one labelled
dataset is divided into a **training set** and a **[blind set](https://yourcloudblog.com/glossary/#blind-evaluation)**. The blind set is held
back, and its answers are sealed away from the model. That split is then **[frozen](https://yourcloudblog.com/glossary/#frozen-experiment-record)**. A
record of what the target means, which rows went where, and which columns may be used is
fixed, and it cannot be edited afterwards. Every later step, from the model build to the
tournament to the certificate, reads that record rather than working things out again.

We keep the roles apart on purpose. **The split decides what the data means. The model
build only does the modelling.** A model never gets to reinterpret its own target or
decide for itself which columns look useful.

### Is any column the answer in disguise?

Each candidate column is examined on the training rows only. The blind rows are never
consulted, so nothing about the test can influence what the model is allowed to see. We
look for:

- **Direct leaks**: a column that reproduces the outcome almost exactly. The bar for
  exclusion is deliberately high. A very strong but legitimate predictor (yesterday's
  price when forecasting today's, for example) must never be thrown out just for being
  good. Such columns go to a person for review instead.
- **Leaks that live in a set of columns**: the commonest real-world leak is not one column
  but several. Maintenance systems record *which kind* of failure happened as separate
  flags, and the flags together *are* the failure. Each one alone looks weak, so any
  one-column-at-a-time check misses it. We search for small groups of low-cardinality
  columns that rebuild the outcome exactly on rows they were not fitted on, and exclude
  the group.
- **Identifiers**: row numbers, customer IDs and account keys let a model memorise
  records instead of learning patterns. They are recognised by their structure, even
  when the split has cut gaps into a running count.
- **Post-event fields**: a column that only gets filled in once the outcome is known,
  such as a call's duration when predicting whether the call will succeed.

One rule matters more than any threshold: **a column's name is evidence, never a
verdict.** A column called `outcome_date` is not excluded because of its name, and a
column with an innocent name is not waved through because of its name. Names can only
support a pattern that has been measured in the data.

### When does the prediction happen?

A value can only be used if it exists at the moment the prediction is requested. So the
experiment has to say what that moment is. It can be a timestamp column, a time built
from several columns, a named business event such as "at booking", or simply *the
moment each observation arrives*. Many datasets have no dates at all, and the last option
covers them. It requires an explicit declaration from the person running the experiment.
It is never assumed, and the declaration becomes part of the frozen record.

### Did every outcome have a fair chance to become known?

Some outcomes take time to appear. A loan defaults months after it is issued. A customer
churns at the end of a contract. If the newest records have not had time to reach their
outcome, they look like successes when they are really just unfinished, and any model
trained or scored on them inherits that bias. We measure how long outcomes take to become
known across the training population and report whether the experiment gave every record
a fair chance.

### Can this blind set support a claim at all?

A blind set can be too small, or contain records that are still open, or be missing an
outcome class entirely. We check that before the answers are opened. We also separate two
things that are often confused. **Reusing** an evaluation population that has been looked
at before is provenance: it is disclosed, and it qualifies the claim. **Contamination**,
where rows from an earlier evaluation have leaked into this training data, is an
integrity failure and blocks certification.

---

## Twelve models, one untouched test

With the experiment frozen, a tournament builds twelve model variants across six
architectures: neural networks of three shapes, random forest, gradient boosting and
logistic regression. All twelve are trained on the same rows and scored against the same
blind set. The winner is picked on balanced measures ([balanced accuracy](https://yourcloudblog.com/glossary/#balanced-accuracy), [macro F1](https://yourcloudblog.com/glossary/#macro-f1)), not
on raw accuracy, which rewards a model for predicting the majority class. The winner can
then be rebuilt on all of the training data. The rebuild is scored on the same held-out
rows, so the two figures are directly comparable.

---

## A certificate is fourteen claims, not one stamp

A single pass/fail verdict hides too much. It cannot tell a model that is sound but
lightly tested from a model that is well tested but was fed a leak. So a certificate is a
set of **independent claims**, each with its own state and a stable reason code that an
auditor can trace:

**Model assurance**: is this the model we say it is?
- **Identity**: the artifact can be hashed and matched exactly.
- **Lineage**: where it came from and what it was built from.
- **Training-data boundary**: what data it was fitted on, and nothing else.
- **Target domain**: which outcomes it was declared to predict.
- **Feature schema**: exactly which inputs it expects.
- **Build reproducibility**: whether rebuilding it produces the same model.
- **Inference repeatability**: whether the same question gets the same answer. This is
  *verified*: the frozen model is asked the same questions twice and the outputs are
  compared. It is never assumed because the algorithm is supposed to be deterministic.

**Performance assurance**: did it earn its score honestly?
- **Observed predictive performance**
- **Evaluation reproducibility**
- **Blind integrity**

**Deployment assurance**: can it be trusted in its real setting?
- **Prediction-time availability**
- **Outcome observability**
- **Evaluation feasibility**

Each claim is one of **CERTIFIED**, **NOT ESTABLISHED**, **NOT ASSESSED**, **NOT
APPLICABLE** or **FAILED**. They are reported side by side and never merged into one.

Two design decisions follow from that:

1. **The model's status comes from the model-assurance claims.** A model whose prediction
   moment was never declared can still be a certified *model*. It is emphatically not a
   certified *deployment*, and the certificate says so in a separate section.
2. **Performance is judged on evidence, not on the verdict.** The performance claim needs
   three things: the metric was actually measured, the evaluation was adequate, and
   training and evaluation were isolated. A strong number from an inadequate test does
   not qualify.

On top of the status, a model earns a **designation**. It is *REPEATABLE* only when the
repeatability check actually ran and passed, and *AUDITABLE* when the evidence is complete
but that check never ran. A failure on any certification gate withholds both.

---

## A metric is only reported where it means something

Suppose a model is supposed to predict three outcomes, but the blind set contains only two
of them. Accuracy can still be computed. Balanced accuracy, macro F1 and per-class recall
cannot honestly be. They would describe a problem the model was never tested on.

So every metric carries a claim of its own. We check the blind set against the **[declared
target domain](https://yourcloudblog.com/glossary/#declared-target-domain)**: the outcomes the model was built to predict, taken from the model itself
or from the frozen experiment. It is never taken from whatever happened to turn up in the
test answers. If a declared class is missing, the evaluation is *inadequate*. Accuracy
stays visible as descriptive, and the class-sensitive metrics read **NOT MEASURABLE**,
with the raw figure shown beside them so nothing is hidden. Coverage and sample size are
reported separately, because neither makes up for the other.

This holds when a model is scored long after it was built, too. A business typically
learns outcomes weeks later, in a new file, against a model whose original data has long
since been cleared away. The model records the outcomes it was fitted on, so it can still
be graded properly from its own artifact alone.

---

## "Not assessed" is an answer

The easiest way to make a certificate look good is to fill every gap with an assumption.
We do the opposite. When the data cannot support a claim, for example when there is no
way to tell when an outcome became known, the claim reads **NOT ASSESSED**, and the
certificate explains why and what would change it.

That is not a failure of the system. It is the system refusing to overstate. A model with
three honest "not assessed" claims is more useful than one with fourteen confident ticks
that nobody can check.

Two more guarantees keep the record trustworthy over time:

- **A certificate that has been issued is never re-judged.** When the methodology
  improves, older models get a *current assessment* rebuilt from their frozen evidence,
  clearly labelled as such, next to the certificate they were originally issued. History
  is not rewritten.
- **The certificate is rebuilt from evidence, not from stored text.** Every claim is
  worked out again from the frozen records, so a certificate cannot drift away from what
  actually happened.

---

## What this looks like on real data

In a recent end-to-end test, an automated agent drove the system on three public
datasets: machine-failure prediction ([AI4I 2020 Predictive Maintenance](https://archive.ics.uci.edu/dataset/601/ai4i+2020+predictive+maintenance+dataset), synthetic, UCI), telecom churn ([IBM Telco Customer Churn](https://www.kaggle.com/datasets/blastchar/telco-customer-churn)) and bank marketing ([UCI Bank Marketing](https://archive.ics.uci.edu/dataset/222/bank+marketing), Moro, Cortez & Rita, 2014). The business results of those trials, in dollars, are in [predictions from business spreadsheets](https://yourcloudblog.com/blog/predictions-from-business-spreadsheets/).

- On the **bank marketing** data, the system flagged *call duration* on its own. That
  column is known only after the call ends, and it is a classic reason published results
  on this dataset look better than they should. It was kept out by default.
- On the **machine-failure** data, it caught the individual failure-mode flags as a
  group. Each flag alone was too weak to trigger a single-column check, but together they
  rebuilt the outcome.
- On the **churn** data, it removed the customer identifier and passed everything else.
- All three models were rebuilt on the full training data and repeated their results
  exactly across three independent runs.

Several verdicts in that test read *not assessed*. Where the data carried no dates, the
experiment had not declared when predictions happen. The certificate said so plainly,
rather than guessing.

---

## What certification is not

It is not a promise of future performance. Markets shift, customers change, and a model
that was honest yesterday can go stale. Certification tells you **the score was earned
honestly, under conditions you can inspect**. It does not tell you the world will stay the
same. Separating those two claims, and being exact about which one is being made, is the
whole point.

A number you can audit beats a better number you cannot.
