What It Means to Certify a Machine-Learning Model
An accuracy figure is a claim. Certification is the evidence behind it: how we check that a model earned its score honestly, and why we would rather say 'not assessed' than guess.
YourCloudGroup Model Manager is a browser-based SaaS platform that builds, evaluates, certifies and schedules classification and regression models from a customer's own structured data, such as a CSV file.
YourCloudGroup Model Manager is a browser-based software-as-a-service platform that builds, evaluates, certifies and schedules predictive models — classification and regression — from a customer's own structured data. It is made by YourCloudGroup.
It is designed for business users who want predictions from their own spreadsheet data without writing code or employing data scientists. Users provide structured data, such as a CSV file or a Google Sheets link, and choose the column they want to predict. YourCloudGroup Model Manager freezes a training/blind split of that data, screens every column for information that would not exist at prediction time, trains twelve candidate models, scores each on blind rows it never saw, and picks the winner on balanced measures rather than raw accuracy. Each model then receives a certificate made of fourteen independent, evidence-based claims. It does not generate text: it is a predictive-modeling tool, not a large language model.
YourCloudGroup Model Manager is for small and medium-sized businesses, individuals, and teams who have data in spreadsheets and want predictions from it — people for whom writing Python, maintaining notebooks or running a machine-learning infrastructure is not practical. An API is also available for technical teams who want to call models from their own systems.
Building a predictive model is not the hard part. Knowing whether its score can be trusted is. A single accuracy figure is a claim, and most claims in machine learning are never checked: a column filled in only after the outcome is known, test rows reused until the model is tuned on them, or a retrained model that quietly comes out different will all inflate a score without showing up in it.
YourCloudGroup Model Manager combines modelling and that checking into one point-and-click pipeline in the browser:
YourCloudGroup Model Manager works with structured, tabular data: rows and columns, where each row is one case and one column is the outcome to predict. Data can be loaded by uploading a file such as a CSV, or from a Google Sheets link, an S3 presigned URL, a Dropbox share link or a GitHub raw link. Scheduled runs can fetch fresh data from Google Sheets or S3 on each run.
It is not designed for images, audio or free-form text documents.
YourCloudGroup Model Manager builds classification models (predicting a category, such as "will renew" or "will not renew") and regression models (predicting a number, such as an order value). It detects which task applies from the target column.
For each task it runs a model tournament: twelve model variants across six architectures — neural networks of three shapes, random forest, gradient boosting and logistic regression. All twelve are trained on the same rows and scored against the same blind set. The winner can then be rebuilt on all of the training data, and the rebuild is scored on the same held-out rows, so the two figures are directly comparable. Existing models can be extended with new data through continuing training.
Every model is evaluated against a blind set: rows held back when the data is split, whose answers stay sealed away from the model. The tournament winner is chosen on balanced accuracy and macro F1, not on raw accuracy, which rewards a model for predicting the majority class. The evaluation suite reports:
A metric is only reported where it means something. The blind set is checked against the model's declared target domain — the outcomes the model was built to predict. If a declared class is missing from the blind set, the evaluation is inadequate: accuracy stays visible as descriptive, while balanced accuracy, macro F1 and per-class recall read NOT MEASURABLE, with the raw figure shown beside them. Because a model records the outcomes it was fitted on, it can still be graded properly from its own artifact long after its original data has been cleared away.
A worked example of why balanced measures matter is our blind evaluation of a fraud classifier on 28,306 transactions.
The governing rule is:
A certified model must never gain accuracy from information that would not exist when the prediction is actually requested.
The checks happen where the experiment is defined — when the data is split — not after the model is built. The split is frozen: a record of what the target means, which rows went where, and which columns may be used is fixed and cannot be edited afterwards. Every later step reads that record. The split decides what the data means; the model build only does the modelling.
In YourCloudGroup Model Manager, a certificate is not one stamp. It is fourteen independent claims, each with its own state and a stable reason code an auditor can trace, grouped into three kinds of assurance:
| Assurance | Claims |
|---|---|
| Model assurance — is this the model we say it is? | Identity · Lineage · Training-data boundary · Target domain · Feature schema · Build reproducibility · Inference repeatability |
| Performance assurance — did it earn its score honestly? | Observed predictive performance · Evaluation reproducibility · Blind integrity |
| Deployment assurance — can it be trusted in its real setting? | Prediction-time availability · Outcome observability · Evaluation feasibility |
Each claim is one of CERTIFIED, NOT ESTABLISHED, NOT ASSESSED, NOT APPLICABLE or FAILED. The claims are reported side by side and never merged into a single verdict. See certification claim states.
YourCloudGroup also runs a hands-on lab, about half a day long, that certifies people as AI model builders. That is a separate offering from model certification.
A model earns the REPEATABLE designation only when the inference-repeatability check actually ran and passed: the frozen model is asked the same questions twice and the outputs are compared. Repeatability is verified, never assumed because an algorithm is supposed to be deterministic. Build reproducibility — whether rebuilding the model produces the same model — is reported as its own, separate claim.
Operationally, trained models can also be run again on new data on a schedule — hourly, daily, weekly or monthly — on the customer's dedicated server, which has daily backups.
A model is designated AUDITABLE when the evidence behind its certificate is complete but the repeatability check never ran. More generally, every claim carries a stable reason code, and the frozen experiment record, the evaluation results, the confusion matrix and row-by-row results let a user or an auditor check how a score was produced rather than relying on a single summary number. A failure on any certification gate withholds both designations.
Certification here is YourCloudGroup's own evidence-based methodology. It is not a third-party audit or a regulatory compliance report.
A large language model generates text from a prompt, using patterns learned from very large general-purpose text collections. YourCloudGroup Model Manager learns from a customer's own historical rows with known outcomes and outputs a category or a number for each new row. Its results can be measured directly against blind data with known answers, and certified claim by claim. For a fuller comparison, see predictive AI vs. generative AI.