# The Model Platform That Told Us the Truth

Canonical: https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/
Published: 2026-09-30
Author: YourCloudGroup · Publisher: YourCloudGroup · Category: Experiments
Tags: case study, model certification, target leakage, blind evaluation, API automation

> An AI agent put YourCloudGroup Model Manager to the test on 77,299 rows of real data. The platform caught six mistakes that would have flattered the result, certified the process, and graded the signal honestly as weak.

**Short answer:** In a 77,299-row experiment run entirely through the API by an AI agent, YourCloudGroup Model Manager caught six problems that would each have inflated the result, certified the process as auditable and repeatable, and graded the model's value as WEAK — 51.6% balanced accuracy against a 50% baseline, a relative lift of 3.2%. Telling the truth before deployment was the most valuable result.
A model platform earns trust by telling you when your model is not good enough. In this case study an AI agent ran a complete, real experiment through [YourCloudGroup Model Manager](https://yourcloudblog.com/model-manager/) on a dedicated YourCloudGroup seat — a private server running the platform — and the platform's most valuable output was an honest one: the process was **CERTIFIED**, and the model's value was graded **WEAK**.

Along the way it caught six mistakes that would each have made the model look better than it was. None of them reached the result.

*This evaluation was carried out by an AI research agent (Claude) acting for a YourCloudGroup user, between 27 and 28 September 2026. Quotations are the agent's. It is not an endorsement by Anthropic. Figures are exact unless marked approximate.*

## The project in one paragraph

An AI agent was asked to find out whether publicly disclosed stock trades by members of the U.S. Congress could predict which stocks would outperform the S&P 500 over the following 60 trading days. It assembled **77,299 disclosures filed between April 2016 and June 2026**, each with 27 facts that were known on the day of disclosure: the trade's size and direction, filing delay, chamber and party, the stock's recent momentum and volatility, and how many other members had traded the same stock recently. The outcome was whether the stock beat SPY. The disclosures are U.S. House and Senate periodic transaction reports filed under the STOCK Act, taken from [QuiverQuant's congressional-trading dataset](https://www.quiverquant.com/congresstrading/); stock and SPY prices are split- and dividend-adjusted daily bars from [Alpaca's market-data API](https://alpaca.markets/). The agent then drove the seat **entirely through its REST API**, with no human in the loop for the modelling: data split, leakage checks, a 12-variant [model tournament](https://yourcloudblog.com/glossary/#model-tournament), certification, full-data rebuild and scoring.

![Bar chart of rows at each stage: 115,812 raw congressional disclosures, 84,065 after cleaning, 77,299 with full price history, 60,698 in the training set (2016 to September 2024), 15,462 in the held-back test (December 2024 to June 2026) and 7,731 graded by the certificate](https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/02_data_pipeline_light.webp)

*From 115,812 raw disclosures to a clean, fair test.*

![Timeline of the chronological split: 60,698 training rows from 2016 to 2024, a 100-day gap, then 15,462 held-back test rows](https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/03_chronological_split_light.webp)

*Train on the past, test on the future: a chronological split with a 100-day gap, so no training outcome overlaps the test.*

## Six mistakes the platform caught before they could flatter the result

A typical AutoML tool would have trained on everything and reported a nicer number. The seat caught these first:

| What went wrong | What the seat did |
|---|---|
| The agent's time-gap setting was misread (days vs. rows) and silently discarded **38,618 rows, half the dataset** | The dry-run split showed exactly how many rows were removed and the training date range that was left. The agent spotted the problem before committing. The seat now also warns automatically when a gap removes more than 25% of the data. |
| The filing date itself was about to be used as a predictor — useless, because every test date lies outside the training range | Flagged as a "time column predictor" leakage risk, with a suggestion to drop it |
| A column looked as if it might describe events *after* the prediction moment | Marked "review required" and excluded by default. The decision went to a human instead of being guessed. |
| A computed column that really did fall after the prediction point | Excluded automatically as post-outcome — a textbook case of [target leakage](https://yourcloudblog.com/glossary/#target-leakage) |
| The test period's prices had drifted from the training period's | Flagged as predictor drift, with the size of the shift, so the result could be read in context |
| The outcome takes 60 days to become known, but that window hadn't been declared | Held back the relevant integrity claim until the window was declared; once it was, the claim became assessed — see [outcome observability](https://yourcloudblog.com/glossary/#outcome-observability) |

![Six cards: half the data about to be lost, a date used as a predictor, an ambiguous timing column, a post-outcome column, an undeclared outcome window, and a shift between periods](https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/04_what_the_seat_caught_light.webp)

*Each of these would have made the model look better than it was.*

> "The most valuable thing the platform did was refuse to let me fool myself. Every one of those issues would have made the model look better than it was." — the AI research agent

## An auditable result, not just a score

After the tournament, the seat produced a **certificate** rather than a single accuracy figure. [Certification in YourCloudGroup Model Manager](https://yourcloudblog.com/blog/what-it-means-to-certify-a-machine-learning-model/) keeps two questions apart:

- **Was the process sound?** It came back **CERTIFIED: AUDITABLE & REPEATABLE PREDICTIVE MODEL**. The integrity claims were all established: prediction-time availability (every input checked against the [prediction moment](https://yourcloudblog.com/glossary/#prediction-moment)), outcome observability, and evaluation feasibility.
- **Does the model add value?** It came back **WEAK, not verified**. The verdict was graded against an explicit baseline ("always predict the most common outcome") with a **90% bootstrap confidence range**: **51.6% [balanced accuracy](https://yourcloudblog.com/glossary/#balanced-accuracy) against 50%** — a **relative lift of 3.2%**, or +1.6 percentage points (90% range 50.7%–52.4%, a relative lift of +1.4% to +4.9%) — on **7,731 held-back rows** the models never saw.
- **Why 7,731 rows.** The seat divides the 15,462 held-back rows in half. One half is used only to pick the winner among the 12 variants; the other half, which played no part in that choice, is the half that is graded — so the verdict is not flattered by the selection.

![Two panels: process CERTIFIED — auditable and repeatable, prediction-time availability, outcome observability and evaluation feasibility established, manifest hash verified, no cross-run contamination; value WEAK — 51.6% balanced accuracy against a 50% baseline, relative lift +3.2% with a 90% range of +1.4% to +4.9%, graded on 7,731 rows that took no part in picking the winner, model value not verified](https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/07_certificate_summary_light.webp)

*Two separate questions, two honest answers.*

The record behind the certificate is complete:

- **Lineage.** Every split has a frozen, hashed manifest (`manifest_hash_verified: true`), and the certificate names the exact training and test files by hash.
- **Test-reuse disclosure.** When the same held-back rows were evaluated again in later runs, the seat recorded it as audit history, and confirmed that none of those rows had leaked into training ("cross-run contamination: not detected"). See [evaluation reuse and contamination](https://yourcloudblog.com/glossary/#evaluation-reuse-and-contamination).

**Why this matters to a business:** when a manager, auditor, regulator or customer asks "how do you know this model works?", there is a document that answers it.

## It told the truth — and that is the feature

Congressional buys and sells beat SPY only about **46–47% of the time**. The best model added a small but real edge overall: **52.7% balanced accuracy against a 50% baseline** after a full-data rebuild, scored on the same graded rows. On closer inspection, its most confident predictions mostly singled out volatile stocks rather than reliable winners.

![Dot chart of balanced accuracy on 7,731 held-back rows: always guessing the common outcome 50.0%; tournament winner on a 10,000-row fit 51.6%, with the certificate's 90% range of 50.7% to 52.4% (relative lift +1.4% to +4.9%); full-data rebuild on all 60,698 rows 52.7%](https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/01_model_vs_baseline_light.webp)

*A small but real edge, graded honestly. 50% balanced accuracy is guessing.*


The seat's grading made this conclusion possible, and it stopped a marginal model from being wired into an automated trading bot.

> "It told us our signal was weak before we deployed it. That's exactly what you want from a model platform: the cheapest model failure is the one you catch before it reaches the business." — the AI research agent

## Built for automation: an AI agent drove it end to end

The agent used only the documented API, and it behaved the way an integration engineer would hope:

- **A readiness check** (`GET /model-types`) confirms reachability, authentication and the available model types in one call.
- **Dry-run splits** (`analyze_only=true`) propose a split, write nothing, and return a `proposal_hash`. The real commit is pinned to that hash, so what gets committed is exactly what was reviewed.
- **Duplicate-safe calls.** Build, train and predict accept an `Idempotency-Key`, so a dropped connection cannot create a duplicate model or charge.
- **Asynchronous tournaments.** Twelve variants across neural networks (basic, deep, wide), logistic regression, random forest and gradient boosting (XGBoost), with polling and progress logs. Each run took about 4–5 minutes on 60,698 training rows.
- **A `/finalise` step** rebuilds the winning design on the *whole* training set and scores it on the *same* held-back rows the tournament graded, so "selected" and "deployable" are compared like for like.
- **Families** give a stable name for a served model, so callers ask for the use case (for example "churn") rather than a model ID, and the business decides when to change which model serves it.
- **Consistent JSON errors** with standard status codes, `Retry-After` on rate limits, and clear refusals rather than guesses — for example `409 ARTIFACT_CHANGED` if a frozen file's hash no longer matches.
- **Readable guides** at `GET /api/v1/docs`, accessible with the API key.

![Pipeline of six steps through one API: dry-run split, commit split pinned to the proposal hash, tournament of 12 variants and 6 model types, certificate with verdict and 90% range and lineage hashes, finalise on all rows, predict and serve with idempotent calls and families](https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/05_api_workflow_light.webp)

*One API, the whole model lifecycle — with no human in the modelling loop.*

## Reported, fixed and redeployed within a day

Heavy use by an automated agent surfaced real issues. What stood out was the turnaround: each was confirmed, fixed, tested and redeployed within a day. Five of the six were live about 9 hours after being reported (overnight, including a wait for the deployment transfer); the sixth took about 12–14 hours.

- **Text inputs silently zeroed at prediction time.** Two high-cardinality text columns reached training but arrived as zeros when scoring. Fixed at the root cause (the prediction path had been re-deciding which columns were free text), with a demonstration test.
- **Row cap on tree-based models.** Random forest and XGBoost builds were limited to 10,000 rows. Now up to 200,000, while tournaments deliberately keep a fair shared sample for comparison.
- **Full-data rebuild** (`/finalise`) added as a first-class endpoint, then corrected about 12–14 hours after it was reported, to carry excluded columns forward.
- **Outcome-evaluation endpoint** (`/predictions/{id}/evaluate`) added, so predictions can be scored against real outcomes using the certificate's own metric definitions.
- **Docs available through the API**, and download links fixed to HTTPS.
- **Name-only timing flags** can now be resolved by a recorded, attributed override, while *measured* timing contradictions still cannot be overridden away.

![Six issues, each marked fixed: text inputs zeroed at scoring, a 10,000-row cap on tree models, no full-data rebuild, rebuild lost excluded columns, no way to score against real outcomes, and guides only behind the web login](https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/06_fixes_shipped_light.webp)

*Five of six fixes were live in about 9 hours; the sixth in about 12–14.*

## Where it fits: predictions you have to stand behind

For predictions on structured business data, a certified conventional model avoids several risks of asking a large language model:

| Need | Certified tabular model | Asking an LLM |
|---|---|---|
| Same input, same answer | Yes: deterministic scoring, verified | Not guaranteed |
| Traceable to the data it learned from | Frozen, hashed training set and manifest | Not traceable |
| Measured accuracy on unseen data | Held-back test, baseline, confidence range | Usually asserted, not measured |
| Explanations | None invented; per-row confidence score | Can produce plausible but unfounded reasoning |
| Data control | Stays on your dedicated seat | Often sent to a third-party model |

Typical business uses are customer churn, loan or invoice default risk, lead scoring, demand or stock-out prediction, claim or fraud triage and maintenance scheduling — cases where a ranked list with confidence scores drives real decisions that may later need justifying. LLMs excel at language and judgement tasks; a certified tabular model excels at measurable prediction on tables. They are complementary; see [predictive AI vs. generative AI](https://yourcloudblog.com/blog/predictive-ai-vs-generative-ai/).

## What this case study shows — and what it does not

**In one sentence:** on 77,299 real, time-ordered records, YourCloudGroup Model Manager caught six problems that would each have inflated the result, certified the process as auditable and repeatable, and graded the model's value honestly as WEAK (51.6% balanced accuracy against a 50% baseline — a relative lift of 3.2%, 90% range 50.7%–52.4%).

It does **not** show that the model, the platform or congressional trades can make money, beat the market, or should drive trading decisions — the finding was the opposite. It is one project, on one dataset, run by one agent over two days; the six fixes describe the platform's state on 27–28 September 2026. The certificate's value verdict applies to the stated held-back rows and does not predict future market behaviour.

## Key facts

| | |
|---|---|
| Dataset | 77,299 rows, 27 inputs, filed April 2016 – June 2026 · sources: QuiverQuant (STOCK Act disclosures), Alpaca adjusted daily prices |
| Split | Chronological: 60,698 training rows (2016 – Sept 2024), 15,462 held-back rows (Dec 2024 – June 2026), 100-day gap |
| Tournament | 12 variants across 6 model types, about 4–5 minutes per run |
| Certificate | CERTIFIED (auditable & repeatable); verdict WEAK; 51.6% vs 50% balanced accuracy, relative lift +3.2% (90% range +1.4% to +4.9%) on 7,731 graded rows; the other 7,731 held-back rows picked the winner |
| Lift units | Relative improvement over the baseline, as a fraction (0.032 = 3.2%). MODERATE starts at 0.2 and STRONG at 0.4, applied to the lower end of the 90% range (here 0.014) |
| Full-data rebuild | 52.7% balanced accuracy vs. 50% baseline, on the same graded rows |
| Issues found and fixed | 6, all within a day (5 in about 9 hours) |

## Takeaway

A score is a claim. The platform worth using is the one that checks the claim before your model touches the business — and says "weak" when the signal is weak. To see how that checking works, read [what it means to certify a machine-learning model](https://yourcloudblog.com/blog/what-it-means-to-certify-a-machine-learning-model/).
