# Do football stats mislead? Five correlations put to the test

Source: https://www.footballdatascience.co.uk/research/correlation-causation
Published: 2026-10-03

> Teams that foul more take fewer points. Sacking the manager is followed by better results. Both are true, and neither means what it seems. Five real Scottish football correlations, re-checked by comparing like with like.

**On the terraces:** Every pundit has a stat that proves their point. This piece takes five that sound convincing and shows what's really behind each one, from fouling to sacking the manager.

## The research question

"Teams that commit more fouls take fewer points." "Sack the manager and results pick up." Both are true. Both are correlations: when one goes up, the other tends to move too. **But does fouling cost points? Does sacking the manager work?** That's a different question, about cause, and the numbers don't answer it on their own.

**Correlation doesn't mean causation** is one of the oldest lines in statistics. This piece takes five real correlations from Scottish football, several of which this site found while checking its own work, and shows the five different ways a correlation can mislead.

## The dataset

Every Scottish Premiership match from 2000/01 to 2025/26 in the football-data.co.uk files, with shots, shots on target, corners, fouls and yellow cards where they were recorded, and half-time scores: **11,754 team-matches** with full stats, grouped into **310 team-seasons** with the stats in at least 30 matches. The manager-change study used **80 in-season changes**, compiled for the [new manager bounce](/myth-or-maths/new-manager-bounce) investigation.

## The method

The idea in every case is the same: **compare like with like.** If two things move together, look for something else that could be moving both, then hold it fixed and see whether the link survives. In practice that means:

- **Holding the hidden cause fixed.** Measure the link only among teams that are alike on the thing that might be driving it.
- **A control group.** Compare what happened with what happened to similar teams that didn't do the thing.
- **Splitting out the giants.** Check whether one or two clubs, far from everyone else, are making the whole pattern.

Two of the five checks are new here; the other three were done in earlier pieces, linked in each section.

## Results

### 1. A hidden cause: does fouling cost points?

Across the 310 team-seasons, physical play, fouls and yellow cards measured against each season's average, has a correlation of **−0.42** with points a game. Teams that foul more take fewer points.

But there's a third thing in the picture: **dominance**, how many more shots, shots on target and corners a team has than it allows. Dominance goes with points (**+0.87**), and dominant teams foul less (**−0.40**): they have the ball, so they rarely need to.

<figure class="rank-chart">
<div role="img" aria-label="Dominance at the top, with arrows down to fewer fouls and to more points. A dashed line between fewer fouls and more points is labelled the correlation we see.">

</div>
<figcaption>Dominance drives both, so fouls and points move together even if fouling itself changed nothing.</figcaption>
</figure>

Hold dominance fixed and the link shrinks from −0.42 to **−0.17**. Split the teams into thirds by dominance and look inside each:

| Teams | Fouling v points |
|---|---|
| Least dominant third | −0.04 |
| Middle third | −0.07 |
| Most dominant third | −0.63 |

For two-thirds of the league, **fouling makes no difference to points at all**. The top third looks different, but it isn't like with like: it runs from Celtic and Rangers down to good seasons from Hearts, Hibs and Aberdeen, and inside it dominance still drives both fouling (−0.63) and points (+0.84). That's the same trap again, one level down. There may be a small real cost to fouling, but it can't be separated from dominance in this data. The [team styles study](/learn/k-means-team-styles) found the same: equally outshot sides took the same points whether they fouled a lot or not.

A third thing that drives both is called a **confounder**, and it's the most common reason a correlation misleads.

### 2. Luck evening out: does sacking the manager work?

Across 80 changes, results did jump after the manager went: from **0.83** points a game in the six games before to **1.39** in the six after. That looks like the new manager working.

<figure class="rank-chart">
<div role="img" aria-label="A run of bad luck at the top, with arrows down to manager sacked and to results recover. A dashed line between them is labelled the correlation we see.">

</div>
<figcaption>Managers are sacked after bad runs, and bad runs are partly luck, which evens out whoever is in charge.</figcaption>
</figure>

But managers are sacked after bad runs, and a bad run is partly bad luck, which tends to even out on its own: **regression to the mean**. The test is a **control group**: sides with the same bad run and the same strength that **kept** their manager. They went from 0.83 to **1.39** too. The extra from changing manager: **0.00**, within −0.13 to +0.13. The full investigation is in [does sacking the manager bring a bounce?](/myth-or-maths/new-manager-bounce)

### 3. A change in the measuring: a skill that wasn't

Some teams score more goals than their shots on target suggest. Does that carry over from season to season, like a finishing skill? At first sight yes: a correlation of **0.42**.

<figure class="rank-chart">
<div role="img" aria-label="How shots were recorded at the top, with arrows down to beating the line this season and beating it next season. A dashed line between them is labelled the correlation we see.">

</div>
<figcaption>The same recording era makes a team beat the line this season and next, so it looks like a skill that carries over.</figcaption>
</figure>

But around 2013/14 the shots counted as on target fell from about 5 a game to about 4.2, while goals hardly moved, so every team after that looks like a better finisher than every team before. Next season is nearly always in the same era as this one. Measure each team against its own season instead and the carry-over falls to **0.24**, and to **0.15** without Celtic and Rangers. Most of the "skill" was the data. The details are in [how many shots is a goal worth?](/learn/linear-regression)

### 4. One group making the pattern: seasons that remember

How well does a club's season match the one five years earlier? Across all clubs, a correlation of **0.71**, as if club strength lasts for years. Take out Celtic and Rangers and it's **0.07**.

Two clubs far above everyone else every season make any season look like any other. For the rest of the league, a season remembers only the last two or three. There's no diagram for this one: it isn't a hidden cause, just two points far from the rest pulling the line. The full story is in [what happens next season?](/learn/time-series-arima)

### 5. Cause running both ways: shots and the score

More shots go with more goals, and more goals go with winning. But the score changes how teams play too: a side that goes ahead can sit back, a side that falls behind has to chase. Cause can run **both ways**.

<figure class="rank-chart">
<div role="img" aria-label="The score and shots, with arrows running both ways between them: shots make goals, and being behind changes how you play.">

</div>
<figcaption>Shots change the score, and the score changes the shots.</figcaption>
</figure>

Can this data show it? Compare each team's full-match shots with its own season average, by the score at half-time:

| At half-time | Team-matches | Shots against own average |
|---|---|---|
| Ahead | 3,567 | +0.77 |
| Level | 4,620 | −0.11 |
| Behind | 3,567 | −0.63 |

Teams behind at half-time ended with **fewer** shots than usual, not more. That doesn't mean chasing teams shoot less. A full-match total mixes the **first half**, where being outshot is partly why a team fell behind, with the **second half**, where the score drives the shooting. With only full-match totals, the two can't be pulled apart, and the honest answer is that **this data can't tell**. It would take shots by half, or minute by minute, which these files don't have. Knowing when your data can't answer a question is part of the job.

## Getting closer to cause

- **Control groups.** Compare with similar teams that didn't do the thing, as in the manager study.
- **Like with like.** Hold the likely hidden cause fixed, and check that the groups really are alike on it.
- **Time order.** A cause comes before its effect. Data that only knows totals can hide which came first.
- **Natural experiments.** Sometimes the world changes one thing for you. In 2020/21 every Scottish match was played behind closed doors: no crowds at all. The home edge in the top two divisions that season was 0.18 goals a game, but one season is too few matches to tell: its range runs from −0.02 to 0.37, wide enough for no home advantage at all or a perfectly normal season. The full analysis is in [is home advantage worth a goal?](/myth-or-maths/home-advantage), and the same season's effect on referees' bookings in [do referees favour the home side?](/research/referees-home-side)
- **Randomised trials.** The gold standard: decide by chance who gets the treatment. Football almost never allows it. No club will toss a coin over sacking its manager.

## Limitations

- **Holding one thing fixed isn't holding everything fixed.** Dominance here is built from shots and corners; money, injuries and the quality of a squad are other possible hidden causes the data doesn't have.
- **Like with like has limits.** Bands of teams are never perfectly alike, as the most dominant third showed.
- **Some questions need better data.** Reverse causation through the score needs shots by half or by minute.
- **No causes are proved here.** Ruling out a false cause is easier than proving a true one.

## Conclusion

All five correlations are real. None means what it first seems: a hidden cause, luck evening out, a change in the recording, two giant clubs and cause running both ways each produced a link that wasn't the effect it looked like. Next time a stat "proves" something, ask what else could move both, and whether the teams being compared are really alike.

## Reproduce the analysis

The results files are published by [football-data.co.uk](https://www.football-data.co.uk/scotlandm.php). Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as `SC0_2425.csv`; they aren't rehosted on this site. This snippet produces the two new checks, the fouling confounder and the half-time shots; the other three come from the snippets in the linked pieces. It runs in a few seconds:

```python
import csv
from collections import defaultdict
from math import sqrt
from statistics import correlation, mean, pstdev

NEED = ("FTHG", "FTAG", "HTHG", "HTAG", "HS", "AS", "HST", "AST", "HC", "AC", "HF", "AF", "HY", "AY")
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

# every match with full stats, seen from each side
sides = []
for s in names:
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        for r in csv.DictReader(f):
            if not all(r.get(c) for c in NEED):
                continue
            v = {c: int(r[c]) for c in NEED}
            for team, us, them in ((r["HomeTeam"], "H", "A"), (r["AwayTeam"], "A", "H")):
                gf, ga = v[f"FT{us}G"], v[f"FT{them}G"]
                ht = v[f"HT{us}G"] - v[f"HT{them}G"]
                sides.append({"season": s, "team": team, "points": 3 if gf > ga else 1 if gf == ga else 0,
                              "half_time": "ahead" if ht > 0 else "level" if ht == 0 else "behind",
                              "x": [v[us + "S"], v[them + "S"], v[us + "ST"], v[them + "ST"], v[us + "C"], v[them + "C"],
                                    v[us + "F"], v[us + "Y"]]})

# 1. a confounder: fouls and cards against points, before and after allowing for dominance
groups = defaultdict(list)
for m in sides:
    groups[m["season"], m["team"]].append(m)
seasons = [{"season": s, "x": [mean(m["x"][j] for m in g) for j in range(8)], "ppg": mean(m["points"] for m in g)}
           for (s, _), g in groups.items() if len(g) >= 30]
for s in names:  # each stat against its own season's league, as in the K-means study
    g = [t for t in seasons if t["season"] == s]
    m, sd = [mean(t["x"][j] for t in g) for j in range(8)], [pstdev([t["x"][j] for t in g]) for j in range(8)]
    for t in g:
        z = [(t["x"][j] - m[j]) / sd[j] for j in range(8)]
        t["dominance"] = (z[0] + z[2] + z[4] - z[1] - z[3] - z[5]) / 6  # shots, on target, corners: for minus against
        t["physical"] = (z[6] + z[7]) / 2                              # fouls and yellow cards
d, p, y = ([t[k] for t in seasons] for k in ("dominance", "physical", "ppg"))
r_py, r_dy, r_dp = correlation(p, y), correlation(d, y), correlation(d, p)
print(f"{len(seasons)} team-seasons")
print(f"physical v points {r_py:+.2f}; dominance v points {r_dy:+.2f}; dominance v physical {r_dp:+.2f}")
partial = (r_py - r_dy * r_dp) / sqrt((1 - r_dy ** 2) * (1 - r_dp ** 2))  # physical v points, dominance held fixed
print(f"physical v points with dominance held fixed: {partial:+.2f}")
ranked = sorted(seasons, key=lambda t: t["dominance"])
for i, label in enumerate(("least dominant third", "middle third", "most dominant third")):
    band = ranked[i * len(ranked) // 3:(i + 1) * len(ranked) // 3]
    print(f"  {label}: physical v points {correlation([t['physical'] for t in band], [t['ppg'] for t in band]):+.2f} ({len(band)} team-seasons)")

top = ranked[2 * len(ranked) // 3:]  # is the top third like with like? not if dominance still drives both inside it
print(f"inside the most dominant third: dominance v physical {correlation([t['dominance'] for t in top], [t['physical'] for t in top]):+.2f}, "
      f"dominance v points {correlation([t['dominance'] for t in top], [t['ppg'] for t in top]):+.2f}")

# 2. cause and effect tangled: full-match shots, against the team's own season average, by the half-time score
for m in sides:
    g = groups[m["season"], m["team"]]
    m["extra_shots"] = m["x"][0] - mean(x["x"][0] for x in g)
for state in ("ahead", "level", "behind"):
    g = [m for m in sides if m["half_time"] == state]
    print(f"{state} at half-time: {len(g)} team-matches, {mean(m['extra_shots'] for m in g):+.2f} shots against their own average")
```

## Further reading

- [Spurious Correlations](https://www.tylervigen.com/spurious-correlations), Tyler Vigen: a collection of real data series that move together for no reason at all, a quick reminder of how easily numbers line up by chance.
- [Causal Inference: What If](https://miguelhernan.org/whatifbook), by Miguel Hernán and James Robins: a free textbook on getting from correlation to cause, including confounding, control groups and natural experiments.
