Glossary of Predictive-Modeling Terms

Plain-language definitions of the predictive-modeling and model-evaluation terms used on the Predictive AI Blog, with YourCloudGroup-specific usage marked explicitly.

These definitions explain the terms used across the Predictive AI Blog. Most are standard machine-learning and statistics terms that YourCloudGroup did not invent. Where YourCloudGroup Model Manager uses a term in a product-specific way, the entry says so.

Predictive model

A predictive model is a mathematical model that learns the relationship between input variables and a known outcome from historical examples, and then estimates that outcome for new cases where it is not yet known.

Predictive models are usually trained on structured data — rows and columns — and produce either a category (classification) or a number (regression). This is a general term.

Classification

Classification is a predictive-modeling task in which the model assigns each case to one of a fixed set of categories, such as "fraud" or "legitimate", or "will renew" or "will not renew".

When there are two categories it is called binary classification; with more, multi-class classification. Classification models are usually measured with a confusion matrix, precision, recall and related metrics.

Regression

Regression is a predictive-modeling task in which the model estimates a numeric value for each case, such as a sale price, an order value or a number of days.

Regression models are usually measured with R-squared, MAE and RMSE.

Target column

The target column is the column in a dataset that holds the outcome a predictive model is trained to predict — for example "churned" or "sale price".

In the historical rows used for training, the target column must contain the real, known outcome. Also called the label, target variable or dependent variable.

Training data

Training data is the portion of a dataset that a predictive model is allowed to learn from; patterns in these rows, including their known outcomes, shape the model.

Performance measured on training data overstates how well a model will do on new cases, which is why models are evaluated on separate held-out data.

Held-out evaluation

Held-out evaluation is the practice of measuring a predictive model on data that was set aside before training and never used to train it, so the result estimates performance on new, unseen cases.

The held-out portion is also called a test set or evaluation set. It is a standard practice in machine learning. YourCloudGroup Model Manager evaluates every model in its tournament against data the model did not train on.

Blind evaluation

A blind evaluation is a held-out evaluation in which the model's predictions are recorded before the actual outcomes of the evaluation rows are revealed and compared.

Locking predictions first removes any opportunity to adjust the model or its predictions after seeing the answers. In YourCloudGroup Model Manager the blind set is fixed when the data is split, recorded in the frozen experiment record, and its answers stay sealed until every prediction is made.

Train/test contamination

Train/test contamination occurs when information from the evaluation data reaches the model during training — through shared or near-duplicate rows, or through repeated tuning against the same test set — making the evaluation no longer independent.

Contamination inflates reported performance in a way that is not visible from the score itself.

Target leakage

Target leakage occurs when a model is trained on an input column that contains information about the outcome that would not be available at the moment a real prediction is made, such as a field recorded only after the outcome happened.

Leakage produces strong test results and weak real-world performance, because the leaked information is missing when predictions are actually needed. Held-out evaluation does not catch it if the leaked column is present in both training and evaluation data.

YourCloudGroup Model Manager screens every candidate column on the training rows only for direct leaks, for leaks spread across a group of columns that together rebuild the outcome, for identifiers and for post-event fields. A column's name is treated as evidence, never as a verdict, and a very strong but legitimate predictor is sent for human review rather than excluded.

Class imbalance

Class imbalance is a situation in classification where one outcome is much rarer than the others — for example, when fraudulent transactions make up less than 1% of all transactions.

Under class imbalance, accuracy can be very high even for a model that never predicts the rare class, so per-class metrics such as recall and balanced accuracy are more informative.

Accuracy

Accuracy is the proportion of all predictions that are correct: correct predictions divided by total predictions.

Accuracy is easy to understand but can be misleading under class imbalance. In one held-out test on 28,306 transactions, a fraud model scored 99.93% accuracy, while a model predicting "legitimate" for every row would have scored 99.82%.

Balanced accuracy

Balanced accuracy is the average of the recall achieved on each class, so every class counts equally regardless of how common it is.

For two classes it is (recall on class A + recall on class B) ÷ 2. In the fraud example above, balanced accuracy was about 86.0%, compared with 99.93% plain accuracy.

Precision

Precision is the proportion of cases a model predicted as a given class that truly belong to that class: true positives divided by all predicted positives.

High precision means few false alarms. In the fraud example, 36 of the 43 transactions flagged as fraud were fraudulent, a precision of 83.7%.

Recall

Recall is the proportion of cases that truly belong to a given class that the model correctly identified: true positives divided by all actual positives.

High recall means few missed cases. Also called sensitivity or true-positive rate. In the fraud example, the model caught 36 of 50 frauds, a recall of 72%.

F1 score

The F1 score is the harmonic mean of precision and recall for a class, a single number between 0 and 1 that is high only when both precision and recall are high.

It is calculated as 2 × precision × recall ÷ (precision + recall).

Macro F1

Macro F1 is the unweighted average of the F1 scores of every class, so rare and common classes contribute equally.

It is useful when the rare class matters as much as the common one, because it is not dominated by the majority class the way accuracy is.

Confusion matrix

A confusion matrix is a table that counts a classification model's predictions against the actual outcomes, showing for each actual class how many cases were predicted as each class.

For two classes it has four cells: true negatives, false positives, false negatives and true positives. Nearly every classification metric can be calculated from it.

R-squared (R²)

R-squared, the coefficient of determination, is a regression metric that measures how much of the variation in the actual values a model's predictions explain, where 1 means perfect prediction and 0 means no better than always predicting the average.

R-squared can be negative on held-out data when a model does worse than predicting the average.

Mean absolute error (MAE)

Mean absolute error is a regression metric equal to the average size of the prediction errors, ignoring their direction, expressed in the same units as the target.

An MAE of 12 on a price target measured in dollars means predictions are off by 12 dollars on average.

Root mean squared error (RMSE)

Root mean squared error is a regression metric equal to the square root of the average squared prediction error, expressed in the same units as the target, which penalises large errors more heavily than MAE does.

When RMSE is much larger than MAE, a few predictions have large errors.

Model tournament

In YourCloudGroup Model Manager, a model tournament is the automatic step that trains twelve model variants across six architectures — neural networks of three shapes, random forest, gradient boosting and logistic regression — on the same training rows, scores all twelve against the same blind set, and picks the winner on balanced accuracy and macro F1 rather than raw accuracy.

The winner can then be rebuilt on all of the training data and scored on the same held-out rows, so the two figures are directly comparable. This is YourCloudGroup Model Manager terminology; the general practice of comparing candidate models on held-out data is usually called model selection.

Reproducibility

Reproducibility is the property that repeating a model-building process with the same data, code and settings produces the same model and the same results.

Full reproducibility requires controlling randomness, software versions and data versions. In YourCloudGroup Model Manager, build reproducibility — whether rebuilding a model produces the same model — is one of the fourteen certification claims, reported with its own state rather than assumed.

Inference repeatability

Inference repeatability is the property that running the same trained model on the same input produces the same prediction every time.

It is distinct from reproducibility of training. In YourCloudGroup Model Manager it is verified, not assumed: the frozen model is asked the same questions twice and the outputs are compared. A model whose check ran and passed can earn the REPEATABLE designation.

Model certification

Model certification is a formal, evidence-based statement that a specific model met defined, checkable criteria — such as being fitted only on its declared data and scored on an adequate, isolated evaluation — at a stated point in time.

In YourCloudGroup Model Manager, a certificate is a set of fourteen independent claims in three groups — model assurance (identity, lineage, training-data boundary, target domain, feature schema, build reproducibility, inference repeatability), performance assurance (observed predictive performance, evaluation reproducibility, blind integrity) and deployment assurance (prediction-time availability, outcome observability, evaluation feasibility) — each with its own state and reason code, never merged into one verdict. Its governing rule: a certified model must never gain accuracy from information that would not exist when the prediction is actually requested.

Certification establishes that a score was earned honestly under conditions that can be inspected. It does not guarantee future performance on data that differs from the evaluation data. YourCloudGroup separately runs a hands-on lab that certifies people as AI model builders.

Certification claim states

In YourCloudGroup Model Manager, every certification claim takes exactly one of five states: CERTIFIED, NOT ESTABLISHED, NOT ASSESSED, NOT APPLICABLE or FAILED.

The states are reported side by side and never combined into a single pass/fail. NOT ASSESSED means the data could not support the claim; the certificate explains why and what would change it, rather than filling the gap with an assumption. This is YourCloudGroup Model Manager terminology.

REPEATABLE and AUDITABLE designations

In YourCloudGroup Model Manager, a model is designated REPEATABLE only when its inference-repeatability check actually ran and passed, and AUDITABLE when its certification evidence is complete but that check never ran.

A failure on any certification gate withholds both designations. This is YourCloudGroup Model Manager terminology.

Frozen experiment record

In YourCloudGroup Model Manager, the frozen experiment record is the fixed, non-editable record made when a dataset is split into training and blind sets: what the target means, which rows went where, and which columns may be used.

Every later step — model build, tournament and certificate — reads this record instead of working things out again, so a model can never reinterpret its own target or choose its own columns. This is YourCloudGroup Model Manager terminology.

Prediction moment

The prediction moment is the point in time at which a prediction is actually requested; only information that exists at that moment may be used as model input.

In YourCloudGroup Model Manager the experiment declares it as a timestamp column, a time built from several columns, a named business event such as "at booking", or the moment each observation arrives. For undated data the last option requires an explicit declaration; it is never assumed.

Outcome observability

Outcome observability is whether every record in a dataset has had enough time for its outcome to become known, so that unfinished records are not mistaken for negative outcomes.

A loan defaults months after it is issued; a customer churns at the end of a contract. YourCloudGroup Model Manager measures how long outcomes take to become known across the training population and reports whether each record had a fair chance, as one of its deployment-assurance claims.

Evaluation adequacy

Evaluation adequacy is whether an evaluation set can support a claim at all — for example, whether it is large enough, free of still-open records, and contains every outcome class the model was built to predict.

In YourCloudGroup Model Manager, adequacy is checked before the blind answers are opened. If a declared class is missing, class-sensitive metrics such as balanced accuracy and macro F1 read NOT MEASURABLE, while accuracy remains visible as descriptive only.

Declared target domain

The declared target domain is the set of outcomes a model was built to predict, taken from the model itself or from the frozen experiment — never inferred from whichever outcomes happen to appear in the test answers.

YourCloudGroup Model Manager checks each blind set against it to decide which metrics can honestly be reported. This is YourCloudGroup Model Manager terminology.

Evaluation reuse and contamination

Evaluation reuse means scoring a model on an evaluation population that has been looked at before; contamination means rows from an earlier evaluation have leaked into the training data.

In YourCloudGroup Model Manager the two are kept distinct: reuse is disclosed provenance that qualifies a claim, while contamination is an integrity failure that blocks certification. See also train/test contamination.

Generative AI

Generative AI refers to models that produce new content — text, images, audio or code — in response to a prompt, rather than predicting a defined outcome for a row of data.

Large language models are the best-known kind of generative AI.

Large language model (LLM)

A large language model is a generative AI model trained on very large collections of text to predict and produce language, used for tasks such as drafting, summarising, answering questions and conversation.

LLMs work with language inputs and outputs. For predicting a defined category or number from structured rows, a predictive model trained on the relevant historical data is the more direct tool; see predictive AI vs. generative AI.