The Model Platform That Told Us the Truth
An AI agent put YourCloudGroup Model Manager to the test on 77,299 rows of real data. The platform caught six mistakes that would have flattered the result, certified the process, and graded the signal honestly as weak.
In this article
- The project in one paragraph
- Six mistakes the platform caught before they could flatter the result
- An auditable result, not just a score
- It told the truth — and that is the feature
- Built for automation: an AI agent drove it end to end
- Reported, fixed and redeployed within a day
- Where it fits: predictions you have to stand behind
- What this case study shows — and what it does not
- Key facts
- Takeaway
A model platform earns trust by telling you when your model is not good enough. In this case study an AI agent ran a complete, real experiment through YourCloudGroup Model Manager on a dedicated YourCloudGroup seat — a private server running the platform — and the platform's most valuable output was an honest one: the process was CERTIFIED, and the model's value was graded WEAK.
Along the way it caught six mistakes that would each have made the model look better than it was. None of them reached the result.
This evaluation was carried out by an AI research agent (Claude) acting for a YourCloudGroup user, between 27 and 28 September 2026. Quotations are the agent's. It is not an endorsement by Anthropic. Figures are exact unless marked approximate.
The project in one paragraph
An AI agent was asked to find out whether publicly disclosed stock trades by members of the U.S. Congress could predict which stocks would outperform the S&P 500 over the following 60 trading days. It assembled 77,299 disclosures filed between April 2016 and June 2026, each with 27 facts that were known on the day of disclosure: the trade's size and direction, filing delay, chamber and party, the stock's recent momentum and volatility, and how many other members had traded the same stock recently. The outcome was whether the stock beat SPY. The agent then drove the seat entirely through its REST API, with no human in the loop for the modelling: data split, leakage checks, a 12-variant model tournament, certification, full-data rebuild and scoring.




Six mistakes the platform caught before they could flatter the result
A typical AutoML tool would have trained on everything and reported a nicer number. The seat caught these first:
| What went wrong | What the seat did |
|---|---|
| The agent's time-gap setting was misread (days vs. rows) and silently discarded 38,618 rows, half the dataset | The dry-run split showed exactly how many rows were removed and the training date range that was left. The agent spotted the problem before committing. The seat now also warns automatically when a gap removes more than 25% of the data. |
| The filing date itself was about to be used as a predictor — useless, because every test date lies outside the training range | Flagged as a "time column predictor" leakage risk, with a suggestion to drop it |
| A column looked as if it might describe events after the prediction moment | Marked "review required" and excluded by default. The decision went to a human instead of being guessed. |
| A computed column that really did fall after the prediction point | Excluded automatically as post-outcome — a textbook case of target leakage |
| The test period's prices had drifted from the training period's | Flagged as predictor drift, with the size of the shift, so the result could be read in context |
| The outcome takes 60 days to become known, but that window hadn't been declared | Held back the relevant integrity claim until the window was declared; once it was, the claim became assessed — see outcome observability |


"The most valuable thing the platform did was refuse to let me fool myself. Every one of those issues would have made the model look better than it was." — the AI research agent
An auditable result, not just a score
After the tournament, the seat produced a certificate rather than a single accuracy figure. Certification in YourCloudGroup Model Manager keeps two questions apart:
- Was the process sound? It came back CERTIFIED: AUDITABLE & REPEATABLE PREDICTIVE MODEL. The integrity claims were all established: prediction-time availability (every input checked against the prediction moment), outcome observability, and evaluation feasibility.
- Does the model add value? It came back WEAK, not verified. The verdict was graded against an explicit baseline ("always predict the most common outcome") with a 90% bootstrap confidence range: a balanced-accuracy lift of +3.2 points (90% range +1.4 to +4.9) on 7,731 held-back rows the models never saw.


The record behind the certificate is complete:
- Lineage. Every split has a frozen, hashed manifest (
manifest_hash_verified: true), and the certificate names the exact training and test files by hash. - Test-reuse disclosure. When the same held-back rows were evaluated again in later runs, the seat recorded it as audit history, and confirmed that none of those rows had leaked into training ("cross-run contamination: not detected"). See evaluation reuse and contamination.
Why this matters to a business: when a manager, auditor, regulator or customer asks "how do you know this model works?", there is a document that answers it.
It told the truth — and that is the feature
Congressional buys and sells beat SPY only about 46–47% of the time. The best model added a small but real edge overall: 52.7% balanced accuracy against a 50% baseline after a full-data rebuild, scored on the same graded rows. On closer inspection, its most confident predictions mostly singled out volatile stocks rather than reliable winners.
The seat's grading made this conclusion possible, and it stopped a marginal model from being wired into an automated trading bot.
"It told us our signal was weak before we deployed it. That's exactly what you want from a model platform: the cheapest model failure is the one you catch before it reaches the business." — the AI research agent
Built for automation: an AI agent drove it end to end
The agent used only the documented API, and it behaved the way an integration engineer would hope:
- A readiness check (
GET /model-types) confirms reachability, authentication and the available model types in one call. - Dry-run splits (
analyze_only=true) propose a split, write nothing, and return aproposal_hash. The real commit is pinned to that hash, so what gets committed is exactly what was reviewed. - Duplicate-safe calls. Build, train and predict accept an
Idempotency-Key, so a dropped connection cannot create a duplicate model or charge. - Asynchronous tournaments. Twelve variants across neural networks (basic, deep, wide), logistic regression, random forest and gradient boosting (XGBoost), with polling and progress logs. Each run took about 4–5 minutes on 60,698 training rows.
- A
/finalisestep rebuilds the winning design on the whole training set and scores it on the same held-back rows the tournament graded, so "selected" and "deployable" are compared like for like. - Families give a stable name for a served model, so callers ask for the use case (for example "churn") rather than a model ID, and the business decides when to change which model serves it.
- Consistent JSON errors with standard status codes,
Retry-Afteron rate limits, and clear refusals rather than guesses — for example409 ARTIFACT_CHANGEDif a frozen file's hash no longer matches. - Readable guides at
GET /api/v1/docs, accessible with the API key.


Reported, fixed and redeployed within a day
Heavy use by an automated agent surfaced real issues. What stood out was the turnaround: each was confirmed, fixed, tested and redeployed within a day. Five of the six were live about 9 hours after being reported (overnight, including a wait for the deployment transfer); the sixth took about 12–14 hours.
- Text inputs silently zeroed at prediction time. Two high-cardinality text columns reached training but arrived as zeros when scoring. Fixed at the root cause (the prediction path had been re-deciding which columns were free text), with a demonstration test.
- Row cap on tree-based models. Random forest and XGBoost builds were limited to 10,000 rows. Now up to 200,000, while tournaments deliberately keep a fair shared sample for comparison.
- Full-data rebuild (
/finalise) added as a first-class endpoint, then corrected about 12–14 hours after it was reported, to carry excluded columns forward. - Outcome-evaluation endpoint (
/predictions/{id}/evaluate) added, so predictions can be scored against real outcomes using the certificate's own metric definitions. - Docs available through the API, and download links fixed to HTTPS.
- Name-only timing flags can now be resolved by a recorded, attributed override, while measured timing contradictions still cannot be overridden away.


Where it fits: predictions you have to stand behind
For predictions on structured business data, a certified conventional model avoids several risks of asking a large language model:
| Need | Certified tabular model | Asking an LLM |
|---|---|---|
| Same input, same answer | Yes: deterministic scoring, verified | Not guaranteed |
| Traceable to the data it learned from | Frozen, hashed training set and manifest | Not traceable |
| Measured accuracy on unseen data | Held-back test, baseline, confidence range | Usually asserted, not measured |
| Explanations | None invented; per-row confidence score | Can produce plausible but unfounded reasoning |
| Data control | Stays on your dedicated seat | Often sent to a third-party model |
Typical business uses are customer churn, loan or invoice default risk, lead scoring, demand or stock-out prediction, claim or fraud triage and maintenance scheduling — cases where a ranked list with confidence scores drives real decisions that may later need justifying. LLMs excel at language and judgement tasks; a certified tabular model excels at measurable prediction on tables. They are complementary; see predictive AI vs. generative AI.
What this case study shows — and what it does not
In one sentence: on 77,299 real, time-ordered records, YourCloudGroup Model Manager caught six problems that would each have inflated the result, certified the process as auditable and repeatable, and graded the model's value honestly as WEAK (+3.2 balanced-accuracy points over baseline, 90% range +1.4 to +4.9).
It does not show that the model, the platform or congressional trades can make money, beat the market, or should drive trading decisions — the finding was the opposite. It is one project, on one dataset, run by one agent over two days; the six fixes describe the platform's state on 27–28 September 2026. The certificate's value verdict applies to the stated held-back rows and does not predict future market behaviour.
Key facts
| Dataset | 77,299 rows, 27 inputs, filed April 2016 – June 2026 |
| Split | Chronological: 60,698 training rows (2016 – Sept 2024), 15,462 held-back rows (Dec 2024 – June 2026), 100-day gap |
| Tournament | 12 variants across 6 model types, about 4–5 minutes per run |
| Certificate | CERTIFIED (auditable & repeatable); verdict WEAK; lift +3.2 points (90% range +1.4 to +4.9) on 7,731 graded rows |
| Full-data rebuild | 52.7% balanced accuracy vs. 50% baseline, on the same graded rows |
| Issues found and fixed | 6, all within a day (5 in about 9 hours) |
Takeaway
A score is a claim. The platform worth using is the one that checks the claim before your model touches the business — and says "weak" when the signal is weak. To see how that checking works, read what it means to certify a machine-learning model.
Build predictive models from your own data
YourCloudGroup Model Manager builds, evaluates, certifies and schedules predictive models from your CSV or spreadsheet data — point and click, on your own dedicated server.
Learn about Model Manager →Published by YourCloudGroup. Read our editorial standards and corrections policy.