Do substitutions work? Changes made while losing, like for like
After a change made while losing, a team's chances rise by almost a quarter and the chances it concedes fall by a third. Compare it with teams in the same spot that hadn't changed yet and most of that happens anyway. 1,696 changes from StatsBomb's open data.
Intermediate
Contents
The research question
On 6 October 2026 Scotland lost 2–1 to Slovenia at Hampden. Ryan Christie put them ahead in the first half, Slovenia equalised just before the hour and went ahead with less than 20 minutes left. Within a couple of minutes Scotland made a double change, Josh McPake on for his debut and Findlay Curtis with him, and pushed for an equaliser that never came. It's the moment every fan has an opinion on: should the changes have come sooner, and were they the right ones?
One match can't answer that. But thousands of matches can answer the question underneath it: when a team is losing and the manager makes a change, does it work? And, the harder part, compared with what?
The dataset
StatsBomb's free open data records every event in each match it covers, including every substitution, with the minute and the reason (tactical or injury), and every shot with its expected goals (xG): how often a chance like that goes in.
Data: StatsBomb open data. The analysis is ours, not StatsBomb's.
Some of the open data follows one club, such as Barcelona's matches across many seasons; that would tell us about one team, so it's left out. What's left is whole seasons and whole tournaments: 3,220 matches. The men's are the 2015/16 Premier League, La Liga, Serie A and Ligue 1, the 2018 and 2022 World Cups, Euro 2020 and 2024, the 2024 Copa América, the 2023 Africa Cup of Nations and the 2021/22 Indian Super League (1,946 matches). The women's are four Women's Super League seasons, the 2023/24 seasons in Spain, Germany and Italy, the 2023 NWSL, the 2019 and 2023 World Cups and the 2022 and 2025 Euros (1,274).
There's no Scottish football in it. Scotland's game is the question, not the evidence.
The method
- The change. A team's first tactical substitution made while losing, between the 60th and 80th minute, when most changes come. Injury changes are left out: they're forced, not chosen. A double change counts once.
- Before and after. The team's chances (non-penalty xG) in the 15 minutes before the change and the 15 minutes after, and the same for the chances it conceded. Both windows sit inside the second half.
- Compared with what? This is the whole question. The comparison is teams at the same minute, losing by the same margin (one goal, or two or more), that hadn't made a change yet, whatever they did later. That's change now against wait.
- Two things a fan would count. Did they score again before the end? Did they get at least a draw?
Every comparison comes with a luck margin, the range that 95% of resamples of the matches fall in.
Results
Taken on its own, a change looks like a masterstroke
There are 1,696 first changes made while losing in the 60th to 80th minute. In the 15 minutes after them, the team's chances rose from 0.82 to 1.01 xG per 90 minutes, up 23%, and the chances it conceded fell from 1.77 to 1.21, down 32%. More going forward, less going wrong. Exactly what a manager hopes a change will do.
Teams that waited picked up too
Now the teams at the same minute and the same score that hadn't changed yet. There were 13,483 such moments, in 1,870 matches.
| xG per 90 | Before | After |
|---|---|---|
| Made, changed | 0.82 | 1.01 |
| Made, waited | 0.82 | 0.89 |
| Conceded, changed | 1.77 | 1.21 |
| Conceded, waited | 1.53 | 1.24 |
Going forward, the two groups started level, and the teams that waited improved too, by 8%: a losing side pushes on whether or not anyone comes off the bench. The teams that changed improved more.
At the back the picture is different. The teams that changed had just been through a worse spell: 1.77 conceded against 1.53. They were more likely to have just let in a goal, too: 45% had conceded in the previous 15 minutes, against 36% of the teams that waited. Managers change things when things have just gone wrong. Bad spells end on their own, and both groups finished at the same level, about 1.2.
That's regression to the mean, and it's the same trap as the new manager bounce: change something at a low point and the recovery that was coming anyway gets the credit. It's also a form of selection bias: the teams that changed weren't picked at random, they picked themselves by having a bad spell. Do football stats mislead? has more traps like it.
Like for like
So the comparison is tightened once more: same minute, same score, and whether they had just conceded. Then the change's own effect is how much more the changers improved than the teams that waited:
| Compared with waiting | Extra effect of the change | Luck margin |
|---|---|---|
| Chances made, xG per 90 | +0.11 | +0.03 to +0.20 |
| Chances conceded, xG per 90 | −0.13 | −0.25 to −0.01 |
| Scored again | 28.1% v 25.8% | +0.1 to +4.5 points |
| Drew or won | 13.4% v 14.0% | −2.1 to +1.0 points |
Going forward, a change does help, and the luck margin stays clear of nothing. But it's small: 0.11 xG per 90 is about 0.02 of a goal over the next 15 minutes. Put another way, losing sides scored again 28.1% of the time after a change and 25.8% of the time when they hadn't changed yet: about one extra goal for every 40 to 45 changes.
At the back, a little of the fall is left once the bad spell is allowed for, only just outside luck. That could be the change, or more of the bad spell ending than "just conceded" catches. This data can't tell them apart.
And in points, nothing shows. Teams losing in the last half hour drew or won about 14% of the time whether they had changed or not, and the luck margin rules out anything bigger than about one extra draw in 100 changes. One goal doesn't always change the result, and one extra goal in 40-odd changes is too rare to show up in 1,696 of them.
The answer was the same in the seasons with three changes allowed (+0.11 per 90, 931 changes) and those from 2020 on, when most competitions allowed five (+0.12, 765).
Limitations
- Which change, not just whether. A striker for a defender isn't the same as like for like. The data says who came on and off, but not what the manager meant by it, so this is the average change, good and bad together.
- Waiting isn't doing nothing. A manager can change the shape, or shout, without a substitution. The teams that waited weren't standing still.
- Not every match is alike. Matching on the minute, the score and a goal just conceded leaves other differences, such as how good the bench is, or whether the opponent changed too.
- Scotland's own game isn't in it. International men's football is a small part of the data, and nothing here says whether those two changes at Hampden were right.
- Not Scottish data. StatsBomb's open data doesn't cover the Scottish leagues; the effect could be different there, though there's no obvious reason it would be.
Conclusion
Changes made while losing work less well than they look. Most of the pick-up after a change, more chances made and fewer conceded, would have come anyway: a losing team pushes on, and managers change things straight after a bad spell, which was going to end. Like for like, a change adds a little going forward, about one extra goal in 40 to 45 changes, and nothing we can measure in points. So the manager who makes a change and sees the side pick up isn't proved right, and the one who waited isn't proved wrong.
Reproduce the analysis
You'll need StatsBomb's open data on your computer, about 15 GB unpacked (the download is compressed). This fetches only the events and the match lists:
Show the Bash3 lines, ready to copy and run.
mkdir -p ~/statsbomb-open-data && curl -sL https://codeload.github.com/statsbomb/open-data/tar.gz/refs/heads/master \
| tar xz -C ~/statsbomb-open-data --strip-components=1 --wildcards 'open-data-master/data/competitions.json' \
'open-data-master/data/matches/*' 'open-data-master/data/events/*'
Then this reads every match once and prints every number in the piece. It takes two or three minutes:
Show the Python120 lines, ready to copy and run.
# From Football Data Science by Bryan McGuire. Free to use with credit.
# https://www.footballdatascience.co.uk/research/do-substitutions-work
import json
import os
import random
from collections import Counter, defaultdict
DATA = os.path.expanduser("~/statsbomb-open-data/data")
WINDOW = 15 # minutes either side of the moment
# Whole seasons and whole tournaments only: skip sets where one team plays in more than a quarter of the matches
games = []
for comp in json.load(open(f"{DATA}/competitions.json")):
matches = json.load(open(f"{DATA}/matches/{comp['competition_id']}/{comp['season_id']}.json"))
seen = Counter(t for m in matches for t in (m["home_team"]["home_team_name"], m["away_team"]["away_team_name"]))
if seen.most_common(1)[0][1] > len(matches) / 4:
continue
for m in matches:
goals, shots, subs = [], [], []
for e in json.load(open(f"{DATA}/events/{m['match_id']}.json")):
if e["period"] > 2: # normal time only
continue
kind, team, minute = e["type"]["name"], e["team"]["name"], e["minute"] + e["second"] / 60
if kind == "Shot" and e["shot"]["outcome"]["name"] == "Goal" or kind == "Own Goal For":
goals.append((e["period"], minute, team))
if kind == "Shot" and e["period"] == 2 and e["shot"]["type"]["name"] != "Penalty":
shots.append((minute, team, e["shot"]["statsbomb_xg"]))
if kind == "Substitution" and e["period"] == 2 and e["substitution"].get("outcome", {}).get("name") == "Tactical":
subs.append((minute, team))
games.append({"teams": (m["home_team"]["home_team_name"], m["away_team"]["away_team_name"]),
"goals": goals, "shots": shots, "subs": subs, "five_subs": comp["season_name"] >= "2020", "women": comp["competition_gender"] == "female"})
print(f"{len(games)} matches: {sum(not g['women'] for g in games)} men's, {sum(g['women'] for g in games)} women's")
def margin(g, team, minute): # goals for minus goals against before this minute of the second half
return sum(1 if t == team else -1 for p, mn, t in g["goals"] if p == 1 or mn < minute)
def xg(g, team, a, b): # non-penalty xG between minutes a and b of the second half
return sum(x for mn, t, x in g["shots"] if t == team and a <= mn < b)
def moment(i, g, team, rival, t): # what the 15 minutes either side of minute t looked like
return {"match": i, "five_subs": g["five_subs"],
"situation": (int(t), max(margin(g, team, t), -2), # the minute, and one goal down or two or more
any(r == rival and p == 2 and t - WINDOW <= mn < t for p, mn, r in g["goals"])), # just conceded?
"for before": xg(g, team, t - WINDOW, t), "for after": xg(g, team, t, t + WINDOW),
"against before": xg(g, rival, t - WINDOW, t), "against after": xg(g, rival, t, t + WINDOW),
"scored again": any(r == team and p == 2 and mn >= t for p, mn, r in g["goals"]),
"drew or won": margin(g, team, 999) >= 0}
changed, waited = [], []
for i, g in enumerate(games):
for team, rival in (g["teams"], g["teams"][::-1]):
subs = [mn for mn, t in g["subs"] if t == team]
# the team's first change while behind, 60th to 80th minute, with no change in the 15 minutes before it
first = next((mn for mn in subs if 60 <= mn < 80 and margin(g, team, mn) < 0), None)
if first is not None and not any(first - WINDOW <= mn < first for mn in subs):
changed.append(moment(i, g, team, rival, first))
# every minute this team was behind and hadn't changed in the last 15 minutes or this minute, whatever it did later
for t in range(60, 80):
if margin(g, team, t) < 0 and not any(t - WINDOW <= mn < t + 1 for mn in subs):
waited.append(moment(i, g, team, rival, t))
def per90(v, field):
return v * 90 / WINDOW if field.startswith(("for", "against")) else v
def compare(ch, wt, field, match_on=3): # each change against teams in the same situation that hadn't changed yet
same = defaultdict(list)
for w in wt:
same[w["situation"][:match_on]].append(w[field])
pairs = [(c[field], sum(s) / len(s)) for c in ch if (s := same.get(c["situation"][:match_on]))]
return (per90(sum(a for a, _ in pairs) / len(pairs), field), per90(sum(b for _, b in pairs) / len(pairs), field), len(pairs))
rnd = random.Random(2026) # fixed, so the luck margins repeat exactly
changed_in, waited_in = defaultdict(list), defaultdict(list)
for c in changed:
changed_in[c["match"]].append(c)
for w in waited:
waited_in[w["match"]].append(w)
def luck_margin(field, before=None, n=500): # 95% range from resampling whole matches
diffs = []
for _ in range(n):
pick = [rnd.randrange(len(games)) for _ in games]
ch = [c for i in pick for c in changed_in[i]]
wt = [w for i in pick for w in waited_in[i]]
a, b, _ = compare(ch, wt, field)
if before:
a0, b0, _ = compare(ch, wt, before)
a, b = a - a0, b - b0
diffs.append(a - b)
diffs.sort()
return diffs[int(0.025 * n)], diffs[int(0.975 * n)]
print(f"first changes while behind, 60th to 80th minute: {len(changed)}; "
f"moments behind without a change: {len(waited)}, in {len({w['match'] for w in waited})} matches")
# 1. Same minute, same margin
print("\nmatched on the minute and the margin. Non-penalty xG per 90, 15 minutes before -> 15 minutes after")
for side in ("for", "against"):
(cb, wb, n), (ca, wa, _) = compare(changed, waited, f"{side} before", 2), compare(changed, waited, f"{side} after", 2)
print(f" {side:8} changed {cb:.2f} -> {ca:.2f} ({ca / cb - 1:+.0%}); waited {wb:.2f} -> {wa:.2f} ({wa / wb - 1:+.0%})")
same = defaultdict(list)
for x in waited:
same[x["situation"][:2]].append(x["situation"][2])
just = [(x["situation"][2], sum(s) / len(s)) for x in changed if (s := same.get(x["situation"][:2]))]
print(f" had conceded in the 15 minutes before: changed {sum(a for a, _ in just) / len(just):.0%}, "
f"waited {sum(b for _, b in just) / len(just):.0%}")
# 2. Like for like: same minute, same margin, and whether they had just conceded
print("\nmatched on the minute, the margin and whether they had just conceded")
for side in ("for", "against"):
(cb, wb, n), (ca, wa, _) = compare(changed, waited, f"{side} before"), compare(changed, waited, f"{side} after")
lo, hi = luck_margin(f"{side} after", f"{side} before")
print(f" xG {side:8} changed {cb:.2f} -> {ca:.2f}; waited {wb:.2f} -> {wa:.2f}; "
f"extra change {(ca - cb) - (wa - wb):+.2f} per 90 ({lo:+.2f} to {hi:+.2f}), {n} changes")
for field in ("scored again", "drew or won"):
c, w, _ = compare(changed, waited, field)
lo, hi = luck_margin(field)
print(f" {field}: changed {c:.1%}, waited {w:.1%}, difference {c - w:+.1%} ({lo:+.1%} to {hi:+.1%})")
for label, era in (("three-sub era", False), ("five-sub era", True)):
ch, wt = [c for c in changed if c["five_subs"] == era], [w for w in waited if w["five_subs"] == era]
(cb, wb, n), (ca, wa, _) = compare(ch, wt, "for before"), compare(ch, wt, "for after")
print(f" {label}: {n} changes, extra change in xG for {(ca - cb) - (wa - wb):+.2f} per 90")
Further reading
- StatsBomb open data, on GitHub (now under Hudl, which owns StatsBomb): the free event data used here, with documentation of every field, including substitutions and the reason for each.