# One good season or a good model? Cross-validation

Source: https://www.footballdatascience.co.uk/learn/cross-validation
Published: 2026-09-29

> A single test season can flatter a model or bury it. Cross-validation tests it again and again on different slices of the data; for football, that means walking forward through the seasons. 23 SPFL seasons show why it matters.

## The football question

We've trained the model, tested it on matches it hasn't seen, and the results look pretty decent. **So are we done?**

No, not quite. What if we just happened to test it on a season that suited the model?

In one season, the team-strength model from [training data and test data](/learn/training-and-test-data) got **57.9%** of match outcomes right. That sounds good. But football never sits still. Managers come and go, players leave, teams go up and down, tactics shift, and even home advantage moves about over time. One season is one set of circumstances.

## The concept

**Cross-validation** means testing a model more than once, on different slices of the data, and looking at all the results together rather than trusting one.

The standard version is **k-fold cross-validation**. Shuffle the matches, cut them into *k* groups (often 5 or 10), then take turns: train on all the groups but one, test on the one left out, and repeat until every group has been the test set once. You get *k* scores instead of one.

For football, I wouldn't just shuffle 23 seasons of matches into random groups. A model predicting 2018/19 shouldn't be learning from games played years later. Time matters, and a shuffle lets the future leak into training, the trap from [training data and test data](/learn/training-and-test-data).

So football uses **walk-forward** validation instead. Train on what we knew at the time, test on what happened next, move the window forward and go again.

<figure class="rank-chart">
<div role="img" aria-label="Two diagrams. K-fold: five rows, each with five blocks of shuffled matches; a different block is the test set in each row. Walk-forward: five rows sliding along a timeline; three seasons of training followed by one test season, moving one season later in each row.">

</div>
<figcaption>K-fold takes turns with shuffled groups, so a test group can sit before matches the model trained on. Walk-forward keeps time in order: three seasons to learn from, the next one to test, then slide along a season.</figcaption>
</figure>

## A football example

The team-strength model rates each side by its points per game over the three previous Scottish Premiership seasons, and picks whichever side rates higher. Walk it forward from 2003/04 to 2025/26: **23 separate tests**, each season predicted using only the three before it.

<figure class="rank-chart">
<div role="img" aria-label="Line chart of accuracy by test season from 2003/04 to 2025/26. Team strength ranges from 38.6% in 2012/13 to 57.9% in 2022/23 and sits above always picking a home win in all but two seasons.">

</div>
<figcaption>Each season's share of results called right. Same model, same method, every season: the score moves around a lot, but it stays above always picking a home win in 21 seasons of 23.</figcaption>
</figure>

| | Team strength | Always a home win |
|---|---|---|
| Average over 23 seasons | **49.9%** | 43.5% |
| Best season | 57.9% | 51.8% |
| Worst season | 38.6% | 38.6% |
| Typical swing (standard deviation) | 4.0 points | 3.2 points |

On average the model beats the home-win baseline by **6.4 percentage points**, so it is adding something. It did so in **21 of the 23 seasons**, which is stronger evidence than any single score.

Its best season was 57.9% (2022/23) and its worst was 38.6% (2012/13): the same approach, with a very different outcome. 2012/13 was the first season after Rangers dropped out of the top flight, and the pecking order the model had learned from the previous three seasons no longer held; that's a likely reason, though not one tested here. The home-win baseline's own worst season, also 38.6%, was 2020/21, played almost entirely behind closed doors (see [home advantage](/myth-or-maths/home-advantage)).

### How much of the swing is just luck?

Some of that season-to-season movement would happen even if the model's true skill never changed. A season is only 228 matches, and as in [training data and test data](/learn/training-and-test-data), luck alone makes an accuracy near 50% wobble by about

$$\text{SE} = \sqrt{\frac{0.5 \times 0.5}{228}} \approx 3.3 \text{ points}$$

<div class="plain" markdown="1">
In plain football

- **228** is the number of matches in a Premiership season.
- **3.3 points** is how much a model's score would typically move from season to season from luck alone, even if nothing about the model or the league changed.
- The model's real swing is **4.0 points**, only a little more. Most of the difference between a 58% season and a 48% season is the luck of which matches went which way.
</div>

That's the real case for cross-validation. One season's score is mostly signal plus a big dose of luck. Average over 23 seasons and the luck largely cancels out: the 6.4-point gap over the baseline is good to about **±1.6 points**. We can trust that number in a way we can't trust 57.9%.

## Why it matters

My first question, when someone tells me their model is 58% accurate, is: **58% on what?** One season? One test set? Or consistently across lots of different periods?

"The model hit 57.9% in one season" and "across 23 seasons it averaged 49.9%, and moved around quite a bit from one year to the next" are very different statements. The second one isn't as flashy, but it tells you far more about the model.

- **One split can flatter or bury a model.** Tested only on 2022/23, team strength looks excellent; tested only on 2012/13, it looks no better than guessing.
- **Consistency is the evidence.** Beating the baseline in 21 seasons of 23 says far more than the best season.
- **In football, keep time in order.** Walk forward, never shuffle: the model must only ever learn from the past.
- **Cross-validation is for choosing, too.** Comparing two models, or picking how complex to make one, should use the average over many folds, not one lucky split.

## Limitations

- **Walk-forward tests are few.** 23 seasons is 23 scores; that's plenty for an average, but the spread is still rough.
- **Old seasons may not look like new ones.** A model that did well in 2005 may matter less for 2027. Some analysts weight recent folds more.
- **The window is a choice.** Three training seasons suits this model; a different window gives a different set of scores.
- **Accuracy may not be the right score at all.** It only asks whether the favourite won; it ignores how confident the model was.

With football, one good Saturday doesn't tell you much. I want to know what happens when Saturday keeps coming round.

## Try it yourself

Pick a simple rule for your league, such as "the side higher in the table wins". Check how often it was right in each of the last five seasons, not just the last one. What's the best season, the worst, and the average? Which would you quote if you were selling the rule, and which is the truth?

## Reproduce the analysis

The results files are published by [football-data.co.uk](https://www.football-data.co.uk/scotlandm.php). Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as `SC0_2425.csv`; they aren't rehosted on this site. Then:

```python
import csv
from collections import Counter
from math import sqrt
from statistics import mean, pstdev, stdev

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        return [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]

def team_strength(train):  # points per game in training; new teams get the average
    pts, games = Counter(), Counter()
    for r in train:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            games[team] += 1
    avg = sum(pts.values()) / sum(games.values())
    ppg = lambda t: pts[t] / games[t] if games[t] else avg
    return lambda r: "H" if ppg(r["HomeTeam"]) >= ppg(r["AwayTeam"]) else "A"

def accuracy(predict, rows):
    return sum(predict(r) == r["FTR"] for r in rows) / len(rows)

# walk forward: train on the three seasons before, test on the next, then move along a season
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
data = {s: season(s) for s in names}
model, home = [], []
for i in range(3, len(names)):
    predict = team_strength(data[names[i - 3]] + data[names[i - 2]] + data[names[i - 1]])
    test = data[names[i]]
    model.append(accuracy(predict, test))
    home.append(accuracy(lambda r: "H", test))
    print(f"20{names[i][:2]}/{names[i][2:]}  {len(test)} matches  team strength {model[-1]:.1%}  home win {home[-1]:.1%}")

gap = [m - h for m, h in zip(model, home)]
for name, x in (("team strength", model), ("home win", home), ("gap", gap)):
    print(f"{name:13} mean {mean(x):.1%}  sd {pstdev(x):.1%}  best {max(x):.1%}  worst {min(x):.1%}")
print("seasons the model beat a home win:", sum(g > 0 for g in gap), "of", len(gap))
print(f"luck alone, 228 matches at 50%: sd {sqrt(0.25 / 228):.1%}")
print(f"margin on the 23-season average gap: ±{1.96 * stdev(gap) / sqrt(len(gap)):.1%}")
```

## Further reading

- [Cross-validation (statistics)](https://en.wikipedia.org/wiki/Cross-validation_(statistics)), Wikipedia. K-fold, leave-one-out and the other variants, and why each is used.
- [Cross-validation: evaluating estimator performance](https://scikit-learn.org/stable/modules/cross_validation.html), scikit-learn. How to do it in code, including `TimeSeriesSplit` for data in time order.
- [Time series cross-validation](https://otexts.com/fpp3/tscv.html), Hyndman and Athanasopoulos, *Forecasting: Principles and Practice*. Time-series cross-validation, the walk-forward idea, with diagrams.
