Do football stats mislead? Five correlations put to the test
Teams that foul more take fewer points. Sacking the manager is followed by better results. Both are true, and neither means what it seems. Five real Scottish football correlations, re-checked by comparing like with like.
Beginner
Contents
- The research question
- The dataset
- The method
- Results
- 1. A hidden cause: does fouling cost points?
- 2. Luck evening out: does sacking the manager work?
- 3. A change in the measuring: a skill that wasn't
- 4. One group making the pattern: seasons that remember
- 5. Cause running both ways: shots and the score
- Getting closer to cause
- Limitations
- Conclusion
- Reproduce the analysis
- Further reading
The research question
"Teams that commit more fouls take fewer points." "Sack the manager and results pick up." Both are true. Both are correlations: when one goes up, the other tends to move too. But does fouling cost points? Does sacking the manager work? That's a different question, about cause, and the numbers don't answer it on their own.
Correlation doesn't mean causation is one of the oldest lines in statistics. This piece takes five real correlations from Scottish football, several of which this site found while checking its own work, and shows the five different ways a correlation can mislead.
The dataset
Every Scottish Premiership match from 2000/01 to 2025/26 in the football-data.co.uk files, with shots, shots on target, corners, fouls and yellow cards where they were recorded, and half-time scores: 11,754 team-matches with full stats, grouped into 310 team-seasons with the stats in at least 30 matches. The manager-change study used 80 in-season changes, compiled for the new manager bounce investigation.
The method
The idea in every case is the same: compare like with like. If two things move together, look for something else that could be moving both, then hold it fixed and see whether the link survives. In practice that means:
- Holding the hidden cause fixed. Measure the link only among teams that are alike on the thing that might be driving it.
- A control group. Compare what happened with what happened to similar teams that didn't do the thing.
- Splitting out the giants. Check whether one or two clubs, far from everyone else, are making the whole pattern.
Two of the five checks are new here; the other three were done in earlier pieces, linked in each section.
Results
1. A hidden cause: does fouling cost points?
Across the 310 team-seasons, physical play, fouls and yellow cards measured against each season's average, has a correlation of −0.42 with points a game. Teams that foul more take fewer points.
But there's a third thing in the picture: dominance, how many more shots, shots on target and corners a team has than it allows. Dominance goes with points (+0.87), and dominant teams foul less (−0.40): they have the ball, so they rarely need to.
Hold dominance fixed and the link shrinks from −0.42 to −0.17. Split the teams into thirds by dominance and look inside each:
| Teams | Fouling v points |
|---|---|
| Least dominant third | −0.04 |
| Middle third | −0.07 |
| Most dominant third | −0.63 |
For two-thirds of the league, fouling makes no difference to points at all. The top third looks different, but it isn't like with like: it runs from Celtic and Rangers down to good seasons from Hearts, Hibs and Aberdeen, and inside it dominance still drives both fouling (−0.63) and points (+0.84). That's the same trap again, one level down. There may be a small real cost to fouling, but it can't be separated from dominance in this data. The team styles study found the same: equally outshot sides took the same points whether they fouled a lot or not.
A third thing that drives both is called a confounder, and it's the most common reason a correlation misleads.
2. Luck evening out: does sacking the manager work?
Across 80 changes, results did jump after the manager went: from 0.83 points a game in the six games before to 1.39 in the six after. That looks like the new manager working.
But managers are sacked after bad runs, and a bad run is partly bad luck, which tends to even out on its own: regression to the mean. The test is a control group: sides with the same bad run and the same strength that kept their manager. They went from 0.83 to 1.39 too. The extra from changing manager: 0.00, within −0.13 to +0.13. The full investigation is in does sacking the manager bring a bounce?
3. A change in the measuring: a skill that wasn't
Some teams score more goals than their shots on target suggest. Does that carry over from season to season, like a finishing skill? At first sight yes: a correlation of 0.42.
But around 2013/14 the shots counted as on target fell from about 5 a game to about 4.2, while goals hardly moved, so every team after that looks like a better finisher than every team before. Next season is nearly always in the same era as this one. Measure each team against its own season instead and the carry-over falls to 0.24, and to 0.15 without Celtic and Rangers. Most of the "skill" was the data. The details are in how many shots is a goal worth?
4. One group making the pattern: seasons that remember
How well does a club's season match the one five years earlier? Across all clubs, a correlation of 0.71, as if club strength lasts for years. Take out Celtic and Rangers and it's 0.07.
Two clubs far above everyone else every season make any season look like any other. For the rest of the league, a season remembers only the last two or three. There's no diagram for this one: it isn't a hidden cause, just two points far from the rest pulling the line. The full story is in what happens next season?
5. Cause running both ways: shots and the score
More shots go with more goals, and more goals go with winning. But the score changes how teams play too: a side that goes ahead can sit back, a side that falls behind has to chase. Cause can run both ways.
Can this data show it? Compare each team's full-match shots with its own season average, by the score at half-time:
| At half-time | Team-matches | Shots against own average |
|---|---|---|
| Ahead | 3,567 | +0.77 |
| Level | 4,620 | −0.11 |
| Behind | 3,567 | −0.63 |
Teams behind at half-time ended with fewer shots than usual, not more. That doesn't mean chasing teams shoot less. A full-match total mixes the first half, where being outshot is partly why a team fell behind, with the second half, where the score drives the shooting. With only full-match totals, the two can't be pulled apart, and the honest answer is that this data can't tell. It would take shots by half, or minute by minute, which these files don't have. Knowing when your data can't answer a question is part of the job.
Getting closer to cause
- Control groups. Compare with similar teams that didn't do the thing, as in the manager study.
- Like with like. Hold the likely hidden cause fixed, and check that the groups really are alike on it.
- Time order. A cause comes before its effect. Data that only knows totals can hide which came first.
- Natural experiments. Sometimes the world changes one thing for you. In 2020/21 every Scottish match was played behind closed doors: no crowds at all. The home edge in the top two divisions that season was 0.18 goals a game, but one season is too few matches to tell: its range runs from −0.02 to 0.37, wide enough for no home advantage at all or a perfectly normal season. The full analysis is in is home advantage worth a goal?, and the same season's effect on referees' bookings in do referees favour the home side?
- Randomised trials. The gold standard: decide by chance who gets the treatment. Football almost never allows it. No club will toss a coin over sacking its manager.
Limitations
- Holding one thing fixed isn't holding everything fixed. Dominance here is built from shots and corners; money, injuries and the quality of a squad are other possible hidden causes the data doesn't have.
- Like with like has limits. Bands of teams are never perfectly alike, as the most dominant third showed.
- Some questions need better data. Reverse causation through the score needs shots by half or by minute.
- No causes are proved here. Ruling out a false cause is easier than proving a true one.
Conclusion
All five correlations are real. None means what it first seems: a hidden cause, luck evening out, a change in the recording, two giant clubs and cause running both ways each produced a link that wasn't the effect it looked like. Next time a stat "proves" something, ask what else could move both, and whether the teams being compared are really alike.
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. This snippet produces the two new checks, the fouling confounder and the half-time shots; the other three come from the snippets in the linked pieces. It runs in a few seconds:
Show the Python59 lines, ready to copy and run.
import csv
from collections import defaultdict
from math import sqrt
from statistics import correlation, mean, pstdev
NEED = ("FTHG", "FTAG", "HTHG", "HTAG", "HS", "AS", "HST", "AST", "HC", "AC", "HF", "AF", "HY", "AY")
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
# every match with full stats, seen from each side
sides = []
for s in names:
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
for r in csv.DictReader(f):
if not all(r.get(c) for c in NEED):
continue
v = {c: int(r[c]) for c in NEED}
for team, us, them in ((r["HomeTeam"], "H", "A"), (r["AwayTeam"], "A", "H")):
gf, ga = v[f"FT{us}G"], v[f"FT{them}G"]
ht = v[f"HT{us}G"] - v[f"HT{them}G"]
sides.append({"season": s, "team": team, "points": 3 if gf > ga else 1 if gf == ga else 0,
"half_time": "ahead" if ht > 0 else "level" if ht == 0 else "behind",
"x": [v[us + "S"], v[them + "S"], v[us + "ST"], v[them + "ST"], v[us + "C"], v[them + "C"],
v[us + "F"], v[us + "Y"]]})
# 1. a confounder: fouls and cards against points, before and after allowing for dominance
groups = defaultdict(list)
for m in sides:
groups[m["season"], m["team"]].append(m)
seasons = [{"season": s, "x": [mean(m["x"][j] for m in g) for j in range(8)], "ppg": mean(m["points"] for m in g)}
for (s, _), g in groups.items() if len(g) >= 30]
for s in names: # each stat against its own season's league, as in the K-means study
g = [t for t in seasons if t["season"] == s]
m, sd = [mean(t["x"][j] for t in g) for j in range(8)], [pstdev([t["x"][j] for t in g]) for j in range(8)]
for t in g:
z = [(t["x"][j] - m[j]) / sd[j] for j in range(8)]
t["dominance"] = (z[0] + z[2] + z[4] - z[1] - z[3] - z[5]) / 6 # shots, on target, corners: for minus against
t["physical"] = (z[6] + z[7]) / 2 # fouls and yellow cards
d, p, y = ([t[k] for t in seasons] for k in ("dominance", "physical", "ppg"))
r_py, r_dy, r_dp = correlation(p, y), correlation(d, y), correlation(d, p)
print(f"{len(seasons)} team-seasons")
print(f"physical v points {r_py:+.2f}; dominance v points {r_dy:+.2f}; dominance v physical {r_dp:+.2f}")
partial = (r_py - r_dy * r_dp) / sqrt((1 - r_dy ** 2) * (1 - r_dp ** 2)) # physical v points, dominance held fixed
print(f"physical v points with dominance held fixed: {partial:+.2f}")
ranked = sorted(seasons, key=lambda t: t["dominance"])
for i, label in enumerate(("least dominant third", "middle third", "most dominant third")):
band = ranked[i * len(ranked) // 3:(i + 1) * len(ranked) // 3]
print(f" {label}: physical v points {correlation([t['physical'] for t in band], [t['ppg'] for t in band]):+.2f} ({len(band)} team-seasons)")
top = ranked[2 * len(ranked) // 3:] # is the top third like with like? not if dominance still drives both inside it
print(f"inside the most dominant third: dominance v physical {correlation([t['dominance'] for t in top], [t['physical'] for t in top]):+.2f}, "
f"dominance v points {correlation([t['dominance'] for t in top], [t['ppg'] for t in top]):+.2f}")
# 2. cause and effect tangled: full-match shots, against the team's own season average, by the half-time score
for m in sides:
g = groups[m["season"], m["team"]]
m["extra_shots"] = m["x"][0] - mean(x["x"][0] for x in g)
for state in ("ahead", "level", "behind"):
g = [m for m in sides if m["half_time"] == state]
print(f"{state} at half-time: {len(g)} team-matches, {mean(m['extra_shots'] for m in g):+.2f} shots against their own average")
Further reading
- Spurious Correlations, Tyler Vigen: a collection of real data series that move together for no reason at all, a quick reminder of how easily numbers line up by chance.
- Causal Inference: What If, by Miguel Hernán and James Robins: a free textbook on getting from correlation to cause, including confounding, control groups and natural experiments.