YourCloudGroup Model Manager for AI Agents

An AI agent can build, test, certify and serve predictive models end to end through the YourCloudGroup Model Manager REST API — with dry-run splits, idempotent calls, sealed blind evaluation and a certificate of 14 evidence-based claims.

YourCloudGroup Model Manager is a predictive-modeling platform with a REST API designed to be driven by AI agents. An agent — for example an LLM-based assistant working for a business user — can load structured data, split it, build and compare twelve model variants, certify the winner, rebuild it on all training data and serve predictions, with no human in the modelling loop. The agent handles language and orchestration; Model Manager produces the predictions, and the evidence that they can be trusted.

This has been done end to end: in a September 2026 evaluation an AI research agent (Claude) drove the whole lifecycle through the API on 77,299 rows of real data, and the platform caught six mistakes that would each have flattered the result. See the model platform that told us the truth.

Why an agent should hand prediction to a model, not generate it

Language models are good at reading requests, planning steps and explaining results. They are not built to produce a measured, repeatable prediction for each row of a table. When an agent needs to answer "which customers will churn?" or "which machines should we inspect first?", the safer design is to call a predictive model trained on the business's own data:

What the agent needs YourCloudGroup Model Manager Asking an LLM to predict
Same input, same answer Yes — deterministic scoring, verified by the repeatability check Not guaranteed
Accuracy measured before use Sealed blind set, explicit baseline, 90% confidence range Usually asserted, not measured
Traceable to its training data Frozen, hashed training set and manifest Not traceable
Honest when the signal is weak Verdicts such as WEAK and claims marked NOT ASSESSED May sound confident regardless
Data control Stays on the customer's dedicated server Often sent to a third-party model

See also predictive AI vs. generative AI.

What "safe" means here

In this context a safe predictive model is one whose behaviour can be checked rather than trusted:

  • It cannot learn from the future. Every candidate column is screened for target leakage — direct leaks, leaks spread across groups of columns, identifiers and post-event fields — and the experiment declares its prediction moment.
  • Its score was earned on data it never saw. The blind set is frozen when the data is split, and half of it is used only for grading, never for choosing the winner.
  • It is certified claim by claim. Each model receives a certificate of fourteen independent claims, each CERTIFIED, NOT ESTABLISHED, NOT ASSESSED, NOT APPLICABLE or FAILED — never merged into one stamp.
  • It gives the same answer twice. A model is designated REPEATABLE only after the frozen model has been asked the same questions twice and the outputs compared.

Certification establishes how a score was earned; it does not guarantee future performance if conditions change.

The agent-facing API, step by step

  1. Readiness check. GET /model-types confirms reachability, authentication and the available model types in one call.
  2. Dry-run split. A split requested with analyze_only=true proposes a training/blind division, writes nothing, and returns a proposal_hash together with how many rows would be used or removed.
  3. Pinned commit. The real split is committed against that proposal_hash, so what is committed is exactly what the agent reviewed. The experiment record is then frozen.
  4. Tournament. An asynchronous model tournament trains twelve variants — neural networks (basic, deep, wide), logistic regression, random forest and gradient boosting — with polling and progress logs.
  5. Certificate. The winner receives its certificate: process claims, a certificate verdict on value against an explicit baseline, lineage hashes and any test-reuse disclosure.
  6. Finalise. /finalise rebuilds the winning design on the whole training set and scores it on the same held-back rows, so "selected" and "deployable" are compared like for like.
  7. Serve through a family. A model family gives a stable name — such as "churn" — so callers ask for the use case rather than a model ID, and the business decides when a new model takes over.
  8. Evaluate against real outcomes. /predictions/{id}/evaluate scores past predictions against outcomes that arrive later, using the certificate's own metric definitions.

Full guides are served at GET /api/v1/docs on each customer's server, readable with the API key.

Built for unattended use

  • Idempotency keys. Build, train and predict accept an Idempotency-Key, so a dropped connection or a retried call cannot create a duplicate model or charge.
  • Clear refusals instead of guesses. Errors are consistent JSON with standard status codes and Retry-After on rate limits. If a frozen file's hash no longer matches, the API refuses with 409 ARTIFACT_CHANGED rather than silently continuing.
  • Warnings an agent can act on. A split whose time gap would discard more than 25% of the data triggers an automatic warning; columns whose timing is ambiguous are excluded by default and marked for review.
  • Private by design. Each subscription runs on its own dedicated server; the data and models stay there.

Good fits for an agent

Customer churn, loan or invoice default risk, lead scoring, demand or stock-out prediction, claim or fraud triage and maintenance scheduling — any recurring decision where a ranked list with per-row confidence scores drives action, and where the decision may later need justifying. The business trials on three public datasets show what that is worth in dollars.

Limitations

  • Structured data only — rows and columns, not images, audio or free text.
  • Some claims need declarations. Without a date per record, outcome timing reads NOT ASSESSED; columns whose timing depends on business meaning need a reviewer's decision.
  • The API reference is per server. The full reference is served on each customer's own server rather than published here.
  • Certification is not a forecast. It certifies how a score was earned on the stated blind set, not future performance.

Learn more