# What Is Target Leakage—and Why Can It Make a Model Look Better Than It Is?

Canonical: https://yourcloudblog.com/blog/what-is-target-leakage/
Published: 2026-10-05
Author: YourCloudGroup · Publisher: YourCloudGroup · Category: Model Integrity
Tags: target leakage, model integrity, data preparation, model evaluation

> Target leakage happens when a model is trained on information that would not be available at prediction time. A measured example: one leaked column lifted ROC AUC from 0.81 to 0.96 on the same 8,238 held-out calls.

**Short answer:** Target leakage is when a training column contains information about the outcome that would not be known when the prediction is made. In a measured test on the public UCI Bank Marketing data, one such column — call duration — lifted held-out ROC AUC from 0.812 to 0.955. The extra score cannot exist in real use.
Target leakage is when a predictive model is trained on information that would not be available at the moment the prediction is actually needed. The leaked column quietly encodes the answer, so the model learns a shortcut instead of a real pattern.

Leakage makes a model look better than it is because the shortcut is present in both the training data and the held-out test data. The test score is excellent. Then the model goes into use, the leaked information does not exist yet, and the extra performance disappears.

**In one sentence, measured:** on 8,238 held-out calls from the public UCI Bank Marketing dataset, adding a single leaked column — the length of the sales call — raised the model's ROC AUC from **0.812 to 0.955**, and the share of subscribers found in the top 20% of the call list from **66.6% to 90.2%**. None of that improvement is available before the call is made.

## A simple example of target leakage

Suppose you want to predict which invoices will be paid more than 30 days late. Your export includes a column called "sent to collections". Invoices are only sent to collections after they are already late, so that column almost perfectly predicts the target — in historical data.

At prediction time, when the invoice has just been issued, "sent to collections" is always empty. The model has learned to depend on a signal that never exists when it matters.

## A measured example: call duration in bank marketing

The [UCI Bank Marketing dataset](https://archive.ics.uci.edu/dataset/222/bank+marketing) records 41,188 phone calls a Portuguese bank made to sell term deposits, with the outcome (did the customer subscribe?) and 20 other columns. One column, `duration`, is the length of the call in seconds. The dataset's own authors warn that it "highly affects the output target", that it "is not known before a call is performed", and that it "should be discarded if the intention is to have a realistic predictive model".

The reason is easy to see in the raw data. The overall subscription rate is 11.3%. Among the 12,795 calls shorter than two minutes it is 1.3%; among the 3,464 calls longer than ten minutes it is 48.6%. Customers who say yes stay on the line — so call length measures the answer, not the customer.

### What we measured

We ran one controlled experiment. Everything was held constant except a single column:

- **Data:** all 41,188 rows of `bank-additional-full.csv`.
- **Split:** 32,950 training rows and 8,238 held-out rows (928 subscribers), stratified, fixed seed — the same rows in both runs.
- **Models:** gradient boosting and logistic regression from scikit-learn 1.5.2, default settings.
- **The only change:** whether `duration` was a feature.

| Model | Columns used | ROC AUC | Balanced accuracy | Subscribers in top 20% of calls |
|---|---|---:|---:|---:|
| Gradient boosting | with `duration` (leaked) | 0.955 | 0.769 | 90.2% (837 of 928) |
| Gradient boosting | without `duration` | 0.812 | 0.624 | 66.6% (618 of 928) |
| Logistic regression | with `duration` (leaked) | 0.942 | 0.707 | 85.7% |
| Logistic regression | without `duration` | 0.801 | 0.603 | 64.9% |

![Bar chart: in the top 20% of 8,238 held-out calls, the model with the leaked call-duration column finds 90.2% of subscribers, the honest model without it finds 66.6%, and random calling finds 20%](https://yourcloudblog.com/blog/what-is-target-leakage/leakage-before-after.svg)

*Same data, same split, same model; one column changed. Random calling finds 20% of subscribers in 20% of the calls.*

With the leaked column, the gradient-boosting model caught 531 of the 928 held-out subscribers at the default threshold; without it, 244. When we shuffled each column in turn to see what the leaky model depended on, shuffling `duration` cut its ROC AUC by 0.25 — six times more than the next most important column (employment variation rate, 0.041).

### What the result means

Both models were evaluated correctly on rows they had never seen. The leaky model's higher score is real arithmetic on historical data — and it is unobtainable in practice, because the call has not happened when you decide whom to call. The honest model's 0.812 is the number that describes what a sales team would actually get. Reporting the 0.955 would overstate the model by the whole gap between "good" and "nearly perfect".

### What the result does not show

This is one dataset, one random split and two common model types. It shows the size of the effect for this column on this data; it does not predict how large leakage will be in your data, which can be smaller or much larger. A random split also ignores changes over time: the case study in [predictions from business spreadsheets](https://yourcloudblog.com/blog/predictions-from-business-spreadsheets/) used a time-ordered split of the same data and found that subscription rates drifted from 11% to 31%, which is a separate problem from leakage. The experiment is a YourCloudGroup analysis of public data, not a Model Manager run; the exact settings are listed under sources below so anyone can reproduce it.

## Common sources of leakage

- **Post-outcome fields.** Anything recorded after the event: call durations, closure reasons, final statuses, refund flags, follow-up dates.
- **Aggregates computed over the whole dataset,** including future rows — for example, a customer's lifetime total calculated after the period you are predicting.
- **Duplicates of the target** under another name.
- **Identifiers that correlate with the outcome** because of how records were created, such as ID ranges assigned to a particular outcome batch.
- **Time mixing,** where rows from the future are used to predict the past.

## Why held-out evaluation does not catch target leakage

A [held-out test](https://yourcloudblog.com/blog/what-is-held-out-model-evaluation/) protects against memorising specific rows, but leakage lives in a column, not a row. If the leaked column is present in both the training rows and the held-out rows, the held-out score inherits the same shortcut. The experiment above is exactly that case: a properly blind evaluation that still reported 0.955.

## How to detect target leakage

- **Ask the timing question for every column:** would this value be known at the moment of prediction?
- **Be suspicious of results that are too good,** especially on problems that are known to be hard.
- **Look at which features the model relies on most.** A single dominant column — like `duration` above, six times more important than anything else — deserves scrutiny.
- **Test on a later time period** than the training data, where leakage often breaks down.

## How YourCloudGroup Model Manager screens for leakage

[YourCloudGroup Model Manager](https://yourcloudblog.com/model-manager/) checks for leakage when the data is split, before any model is trained. Every candidate column is examined on the training rows only — never the blind rows — for four patterns:

- **Direct leaks:** a column that reproduces the outcome almost exactly. The bar for exclusion is deliberately high, so a very strong but legitimate predictor is sent to a person for review rather than thrown out for being good.
- **Leaks spread across a group of columns:** separate flags that each look weak but together rebuild the outcome — for example, maintenance systems that record which kind of failure happened as several columns.
- **Identifiers:** row numbers, customer IDs and account keys that let a model memorise records instead of learning patterns.
- **Post-event fields:** columns filled in only once the outcome is known.

A column's name is treated as evidence, never as a verdict. The experiment also declares its [prediction moment](https://yourcloudblog.com/glossary/#prediction-moment), because a value can only be used if it exists when the prediction is requested.

In end-to-end tests run through the platform, this screening:

- flagged the same bank-marketing `duration` column as post-event and kept it out by default;
- caught four machine-failure-mode flags (TWF, HDF, PWF, OSF) as a group, even though each flag alone was too weak for a single-column check, and kept the UDI row number out as an identifier;
- removed the customer identifier from a telecom churn dataset and passed every other column;
- in a separate stock-disclosure project, flagged the filing date as a "time column predictor" and sent a column that might describe events after the prediction moment to a person for review.

These detection results come from the [business spreadsheet trials](https://yourcloudblog.com/blog/predictions-from-business-spreadsheets/) and the case study [the model platform that told us the truth](https://yourcloudblog.com/blog/the-model-platform-that-told-us-the-truth/). Leakage screening is also one of the evidence-based claims on a [Model Manager certificate](https://yourcloudblog.com/blog/what-it-means-to-certify-a-machine-learning-model/).

## Limitations

Automated screening reduces leakage risk; it does not eliminate it. A column can leak in ways that statistics on the training rows cannot reveal — for example, a field that a downstream system back-fills weeks later. That is why borderline columns go to a person who knows how the data was recorded, and why the prediction moment must be declared rather than guessed.

## Takeaway

Leakage is not a modelling bug; it is a data question. In the measured example, one column that only exists after the call turned a 0.81 model into an apparent 0.96. Before training, check every column against the moment of prediction, and remove anything that would not yet exist.

## Sources

- Moro, S., Cortez, P. and Rita, P. (2014). *A data-driven approach to predict the success of bank telemarketing.* Decision Support Systems 62, 22–31. Dataset: [UCI Machine Learning Repository — Bank Marketing](https://archive.ics.uci.edu/dataset/222/bank+marketing), licensed CC BY 4.0.
- YourCloudGroup experiment, 5 October 2026: stratified 80/20 split with seed 42, scikit-learn 1.5.2, input file md5 `f6cb2c1256ffe2836b36df321f46e92c`, `HistGradientBoostingClassifier(random_state=42)` and `LogisticRegression(max_iter=2000)` with one-hot encoded categories, default 0.5 threshold for balanced accuracy.
