The Model Platform That Told Us the Truth

An AI agent put YourCloudGroup Model Manager to the test on 77,299 rows of real data. The platform caught six mistakes that would have flattered the result, certified the process, and graded the signal honestly as weak.

Case study cover: The Model Platform That Told Us the Truth — process CERTIFIED, verdict WEAK, six issues fixed within a day Case study cover: The Model Platform That Told Us the Truth — process CERTIFIED, verdict WEAK, six issues fixed within a day
In this article
  1. The project in one paragraph
  2. Six mistakes the platform caught before they could flatter the result
  3. An auditable result, not just a score
  4. It told the truth — and that is the feature
  5. Built for automation: an AI agent drove it end to end
  6. Reported, fixed and redeployed within a day
  7. Where it fits: predictions you have to stand behind
  8. What this case study shows — and what it does not
  9. Key facts
  10. Takeaway

A model platform earns trust by telling you when your model is not good enough. In this case study an AI agent ran a complete, real experiment through YourCloudGroup Model Manager on a dedicated YourCloudGroup seat — a private server running the platform — and the platform's most valuable output was an honest one: the process was CERTIFIED, and the model's value was graded WEAK.

Along the way it caught six mistakes that would each have made the model look better than it was. None of them reached the result.

This evaluation was carried out by an AI research agent (Claude) acting for a YourCloudGroup user, between 27 and 28 September 2026. Quotations are the agent's. It is not an endorsement by Anthropic. Figures are exact unless marked approximate.

The project in one paragraph

An AI agent was asked to find out whether publicly disclosed stock trades by members of the U.S. Congress could predict which stocks would outperform the S&P 500 over the following 60 trading days. It assembled 77,299 disclosures filed between April 2016 and June 2026, each with 27 facts that were known on the day of disclosure: the trade's size and direction, filing delay, chamber and party, the stock's recent momentum and volatility, and how many other members had traded the same stock recently. The outcome was whether the stock beat SPY. The agent then drove the seat entirely through its REST API, with no human in the loop for the modelling: data split, leakage checks, a 12-variant model tournament, certification, full-data rebuild and scoring.

Bar chart of rows at each stage: 115,812 raw congressional disclosures, 84,065 after cleaning, 77,299 with full price history, 60,698 in the training set (2016 to September 2024), 15,462 in the held-back test (December 2024 to June 2026) and 7,731 graded by the certificateBar chart of rows at each stage: 115,812 raw congressional disclosures, 84,065 after cleaning, 77,299 with full price history, 60,698 in the training set (2016 to September 2024), 15,462 in the held-back test (December 2024 to June 2026) and 7,731 graded by the certificate
From 115,812 raw disclosures to a clean, fair test.
Timeline of the chronological split: 60,698 training rows from 2016 to 2024, a 100-day gap, then 15,462 held-back test rowsTimeline of the chronological split: 60,698 training rows from 2016 to 2024, a 100-day gap, then 15,462 held-back test rows
Train on the past, test on the future: a chronological split with a 100-day gap, so no training outcome overlaps the test.

Six mistakes the platform caught before they could flatter the result

A typical AutoML tool would have trained on everything and reported a nicer number. The seat caught these first:

What went wrong What the seat did
The agent's time-gap setting was misread (days vs. rows) and silently discarded 38,618 rows, half the dataset The dry-run split showed exactly how many rows were removed and the training date range that was left. The agent spotted the problem before committing. The seat now also warns automatically when a gap removes more than 25% of the data.
The filing date itself was about to be used as a predictor — useless, because every test date lies outside the training range Flagged as a "time column predictor" leakage risk, with a suggestion to drop it
A column looked as if it might describe events after the prediction moment Marked "review required" and excluded by default. The decision went to a human instead of being guessed.
A computed column that really did fall after the prediction point Excluded automatically as post-outcome — a textbook case of target leakage
The test period's prices had drifted from the training period's Flagged as predictor drift, with the size of the shift, so the result could be read in context
The outcome takes 60 days to become known, but that window hadn't been declared Held back the relevant integrity claim until the window was declared; once it was, the claim became assessed — see outcome observability
Six cards: half the data about to be lost, a date used as a predictor, an ambiguous timing column, a post-outcome column, an undeclared outcome window, and a shift between periodsSix cards: half the data about to be lost, a date used as a predictor, an ambiguous timing column, a post-outcome column, an undeclared outcome window, and a shift between periods
Each of these would have made the model look better than it was.

"The most valuable thing the platform did was refuse to let me fool myself. Every one of those issues would have made the model look better than it was." — the AI research agent

An auditable result, not just a score

After the tournament, the seat produced a certificate rather than a single accuracy figure. Certification in YourCloudGroup Model Manager keeps two questions apart:

  • Was the process sound? It came back CERTIFIED: AUDITABLE & REPEATABLE PREDICTIVE MODEL. The integrity claims were all established: prediction-time availability (every input checked against the prediction moment), outcome observability, and evaluation feasibility.
  • Does the model add value? It came back WEAK, not verified. The verdict was graded against an explicit baseline ("always predict the most common outcome") with a 90% bootstrap confidence range: a balanced-accuracy lift of +3.2 points (90% range +1.4 to +4.9) on 7,731 held-back rows the models never saw.
Two panels: process CERTIFIED — auditable and repeatable, prediction-time availability, outcome observability and evaluation feasibility established, manifest hash verified, no cross-run contamination; value WEAK — +3.2 points balanced accuracy over baseline, 90% range +1.4 to +4.9 points, 7,731 held-back rows, model value not verifiedTwo panels: process CERTIFIED — auditable and repeatable, prediction-time availability, outcome observability and evaluation feasibility established, manifest hash verified, no cross-run contamination; value WEAK — +3.2 points balanced accuracy over baseline, 90% range +1.4 to +4.9 points, 7,731 held-back rows, model value not verified
Two separate questions, two honest answers.

The record behind the certificate is complete:

  • Lineage. Every split has a frozen, hashed manifest (manifest_hash_verified: true), and the certificate names the exact training and test files by hash.
  • Test-reuse disclosure. When the same held-back rows were evaluated again in later runs, the seat recorded it as audit history, and confirmed that none of those rows had leaked into training ("cross-run contamination: not detected"). See evaluation reuse and contamination.

Why this matters to a business: when a manager, auditor, regulator or customer asks "how do you know this model works?", there is a document that answers it.

It told the truth — and that is the feature

Congressional buys and sells beat SPY only about 46–47% of the time. The best model added a small but real edge overall: 52.7% balanced accuracy against a 50% baseline after a full-data rebuild, scored on the same graded rows. On closer inspection, its most confident predictions mostly singled out volatile stocks rather than reliable winners.

The seat's grading made this conclusion possible, and it stopped a marginal model from being wired into an automated trading bot.

"It told us our signal was weak before we deployed it. That's exactly what you want from a model platform: the cheapest model failure is the one you catch before it reaches the business." — the AI research agent

Built for automation: an AI agent drove it end to end

The agent used only the documented API, and it behaved the way an integration engineer would hope:

  • A readiness check (GET /model-types) confirms reachability, authentication and the available model types in one call.
  • Dry-run splits (analyze_only=true) propose a split, write nothing, and return a proposal_hash. The real commit is pinned to that hash, so what gets committed is exactly what was reviewed.
  • Duplicate-safe calls. Build, train and predict accept an Idempotency-Key, so a dropped connection cannot create a duplicate model or charge.
  • Asynchronous tournaments. Twelve variants across neural networks (basic, deep, wide), logistic regression, random forest and gradient boosting (XGBoost), with polling and progress logs. Each run took about 4–5 minutes on 60,698 training rows.
  • A /finalise step rebuilds the winning design on the whole training set and scores it on the same held-back rows the tournament graded, so "selected" and "deployable" are compared like for like.
  • Families give a stable name for a served model, so callers ask for the use case (for example "churn") rather than a model ID, and the business decides when to change which model serves it.
  • Consistent JSON errors with standard status codes, Retry-After on rate limits, and clear refusals rather than guesses — for example 409 ARTIFACT_CHANGED if a frozen file's hash no longer matches.
  • Readable guides at GET /api/v1/docs, accessible with the API key.
Pipeline of six steps through one API: dry-run split, commit split pinned to the proposal hash, tournament of 12 variants and 6 model types, certificate with verdict and 90% range and lineage hashes, finalise on all rows, predict and serve with idempotent calls and familiesPipeline of six steps through one API: dry-run split, commit split pinned to the proposal hash, tournament of 12 variants and 6 model types, certificate with verdict and 90% range and lineage hashes, finalise on all rows, predict and serve with idempotent calls and families
One API, the whole model lifecycle — with no human in the modelling loop.

Reported, fixed and redeployed within a day

Heavy use by an automated agent surfaced real issues. What stood out was the turnaround: each was confirmed, fixed, tested and redeployed within a day. Five of the six were live about 9 hours after being reported (overnight, including a wait for the deployment transfer); the sixth took about 12–14 hours.

  • Text inputs silently zeroed at prediction time. Two high-cardinality text columns reached training but arrived as zeros when scoring. Fixed at the root cause (the prediction path had been re-deciding which columns were free text), with a demonstration test.
  • Row cap on tree-based models. Random forest and XGBoost builds were limited to 10,000 rows. Now up to 200,000, while tournaments deliberately keep a fair shared sample for comparison.
  • Full-data rebuild (/finalise) added as a first-class endpoint, then corrected about 12–14 hours after it was reported, to carry excluded columns forward.
  • Outcome-evaluation endpoint (/predictions/{id}/evaluate) added, so predictions can be scored against real outcomes using the certificate's own metric definitions.
  • Docs available through the API, and download links fixed to HTTPS.
  • Name-only timing flags can now be resolved by a recorded, attributed override, while measured timing contradictions still cannot be overridden away.
Six issues, each marked fixed: text inputs zeroed at scoring, a 10,000-row cap on tree models, no full-data rebuild, rebuild lost excluded columns, no way to score against real outcomes, and guides only behind the web loginSix issues, each marked fixed: text inputs zeroed at scoring, a 10,000-row cap on tree models, no full-data rebuild, rebuild lost excluded columns, no way to score against real outcomes, and guides only behind the web login
Five of six fixes were live in about 9 hours; the sixth in about 12–14.

Where it fits: predictions you have to stand behind

For predictions on structured business data, a certified conventional model avoids several risks of asking a large language model:

Need Certified tabular model Asking an LLM
Same input, same answer Yes: deterministic scoring, verified Not guaranteed
Traceable to the data it learned from Frozen, hashed training set and manifest Not traceable
Measured accuracy on unseen data Held-back test, baseline, confidence range Usually asserted, not measured
Explanations None invented; per-row confidence score Can produce plausible but unfounded reasoning
Data control Stays on your dedicated seat Often sent to a third-party model

Typical business uses are customer churn, loan or invoice default risk, lead scoring, demand or stock-out prediction, claim or fraud triage and maintenance scheduling — cases where a ranked list with confidence scores drives real decisions that may later need justifying. LLMs excel at language and judgement tasks; a certified tabular model excels at measurable prediction on tables. They are complementary; see predictive AI vs. generative AI.

What this case study shows — and what it does not

In one sentence: on 77,299 real, time-ordered records, YourCloudGroup Model Manager caught six problems that would each have inflated the result, certified the process as auditable and repeatable, and graded the model's value honestly as WEAK (+3.2 balanced-accuracy points over baseline, 90% range +1.4 to +4.9).

It does not show that the model, the platform or congressional trades can make money, beat the market, or should drive trading decisions — the finding was the opposite. It is one project, on one dataset, run by one agent over two days; the six fixes describe the platform's state on 27–28 September 2026. The certificate's value verdict applies to the stated held-back rows and does not predict future market behaviour.

Key facts

Dataset 77,299 rows, 27 inputs, filed April 2016 – June 2026
Split Chronological: 60,698 training rows (2016 – Sept 2024), 15,462 held-back rows (Dec 2024 – June 2026), 100-day gap
Tournament 12 variants across 6 model types, about 4–5 minutes per run
Certificate CERTIFIED (auditable & repeatable); verdict WEAK; lift +3.2 points (90% range +1.4 to +4.9) on 7,731 graded rows
Full-data rebuild 52.7% balanced accuracy vs. 50% baseline, on the same graded rows
Issues found and fixed 6, all within a day (5 in about 9 hours)

Takeaway

A score is a claim. The platform worth using is the one that checks the claim before your model touches the business — and says "weak" when the signal is weak. To see how that checking works, read what it means to certify a machine-learning model.

YourCloudGroup

Build predictive models from your own data

YourCloudGroup Model Manager builds, evaluates, certifies and schedules predictive models from your CSV or spreadsheet data — point and click, on your own dedicated server.

Learn about Model Manager →

Published by YourCloudGroup. Read our editorial standards and corrections policy.