Expected goals from scratch. Where everyone stood
Most expected-goals models only know where a shot was taken. Train one on 95,162 real shots that also record where the keeper and defenders stood, and it gets within touching distance of a professional model. Then drag the players yourself.
Intermediate Part 20 of Machine Learning Through Football
New to the notation? The symbols explained
Contents
The football question
A striker shoots from the edge of the box and it's saved. Was that a good chance? Expected goals, or xG, answers with a number: how often a shot like that goes in. An xG of 0.10 means about one in ten.
Most xG you see on television starts from where the shot was taken. But you know from watching that the same spot can be a great chance or a hopeless one. It depends on who's in the way and where the keeper is. What if the model could see that too?
This part builds one from scratch, on real shots that record where everyone stood, and tests it on matches it never saw. Then you can drag the players around yourself in the Shot Lab.
The concept
It's the logistic regression from part 8, with one outcome instead of three: goal or no goal. Each thing the model knows about a shot is a feature. Each feature gets a weight, the weighted features are added up, and the total is squeezed into a probability between 0 and 1:
$$z = b + w_1 x_1 + w_2 x_2 + \dots$$
$$p = \frac{1}{1 + e^{-z}}$$
In plain football
- x₁, x₂, … are the features: how far out, how much of the goal you can see, how many defenders are in the way, and so on.
- w₁, w₂, … are the weights the model learns. A negative weight means more of that feature makes a goal less likely.
- b is where every shot starts before the features push it up or down.
- p is the xG: the chance this shot goes in. The second line just turns any total, however big or small, into a chance between 0 and 1.
The weights are found by gradient descent: start at zero, check how wrong the chances are on the training shots, nudge every weight the way that makes them less wrong, and repeat a thousand times. Wrong is measured by log loss, the score from model evaluation: it punishes a confident miss hardest. Lower is better.
What a freeze frame adds
The data is StatsBomb's free open data: every event in 3,961 matches, from recent men's and women's World Cups and Euros to the 2015/16 season in England, Spain, Italy and France and four Women's Super League seasons. For each shot it records a freeze frame: where the keeper and every nearby player stood at the moment the ball was struck.
Data: StatsBomb open data. The analysis is ours, not StatsBomb's.
Leaving out penalties, free kicks and corners, which are different problems, leaves 95,162 open-play shots with the keeper in the frame. 10.4% of them went in. About a third, 35%, come from women's matches.
Here's one made-up shot from the edge of the box, before and after. Same spot, same keeper; on the right, the three defenders have stepped out of the pale triangle between the ball and the posts:
xG 0.07
xG 0.10
Knowing only the distance, the model gives both shots 0.07. Knowing where the defenders stood, it gives the open one 0.10: nearly half as likely again to go in.
Measuring a shot
The positions become nine features, all measured from the freeze frame:
- Distance to the centre of the goal, in metres. It goes in twice, as it is and as a logarithm, which lets the model draw a curve rather than a straight line.
- Angle: how much of the goal mouth the shooter can see, between the two posts.
- Header: head or foot.
- Defenders in the way: how many outfield opponents stand inside the triangle from the ball to the posts.
- Nearest defender: how far away the closest one is, counted up to 3 metres.
- Keeper in the way: is the keeper inside the triangle?
- Keeper off-centre: how far the keeper is from the straight line between ball and goal.
- Keeper off the line: how far out of goal the keeper has come.
- Goal you can see: how much of the goal mouth is left once every body in front of the ball is taken away, each blocking half a yard either side. It's the angle, minus whatever the keeper and defenders hide.
The matches were split at random: 57,154 shots to train on, 18,951 held back for making choices, and 19,057 kept back for one final test. Splitting by match, not by shot, keeps every shot from one game on the same side, as in training and test data. Choices such as the 3-metre cap, the two forms of distance and the half-yard body width were made on the held-back shots; the test shots were used once, at the end.
Three models, one test
Three models, each knowing more than the last, refitted on the training and held-back shots together, then scored on the test shots:
| Model | Log loss | Brier |
|---|---|---|
| Guess 10.4% every time | 0.333 | 0.093 |
| Distance only | 0.297 | 0.086 |
| Distance and angle | 0.294 | 0.084 |
| Everything | 0.266 | 0.076 |
| StatsBomb's own xG | 0.263 | 0.075 |
Distance does most of the early work. The angle adds a little, because distance and angle mostly tell the same story. The freeze frame adds far more than the angle does: from 0.294 to 0.266. On both scores the full model gets more than nine-tenths of the way from guessing to StatsBomb's professional model.
What it learned
The weights are per standard deviation of each feature, so they can be compared:
- Distance dominates, −0.96 as it is and +0.04 as a logarithm: further out, fewer goals, at an almost steady rate.
- Goal you can see, +0.39, and the plain angle, +0.11: once the model can see how much of the goal is actually open, the angle between the posts has much less left to say.
- Defenders in the way, −0.16, and space, +0.20: bodies near the ball or in the line cost chances, on top of the goal they hide.
- Header, −0.30. From six yards out with nobody in the way, a shot with the foot is 0.61, a header 0.41.
- Keeper in the way, −0.05, and keeper off-centre, −0.03: close to nothing, once the model knows how much goal is left to aim at.
- Keeper off the line, +0.38: the further out the keeper, the more likely a goal. A one-on-one with the keeper rushing out comes out at 0.34; the same shot with the keeper on the line, 0.23.
That last one needs care. Keepers come out to narrow the angle, and coaches teach it for good reason. But in the data, a keeper far off the line is usually a keeper in trouble: rounded, caught by a through ball, or scrambling after a rebound. The model learns that keepers off their line concede more; it can't tell whether coming out caused it. It's the trap in do football stats mislead?: a pattern is not a cause until you've compared like with like.
Could the model be missing something that would explain it away? The obvious suspect is how much goal the keeper covers, so the goal you can see went in partly to test that. It didn't explain it: the keeper-off-the-line weight stayed just as big. And taking the feature out makes the forecasts on the held-back shots clearly worse, 0.2758 against 0.2691, so it stays in: it carries real information about keepers in trouble, even though it isn't advice.
Men's and women's shots
Do women's and men's shots need different models? The test for it is simple: give the model one more feature, "this is a women's match", and see whether it helps on the held-back shots. It made no real difference: log loss 0.2690 with it and 0.2691 without. Once you know where everyone stood, a shot is a shot. So one model serves both.
Against the professionals
StatsBomb's own xG is in the same files, so it can be scored on the same test shots: 0.263 against our 0.266. That small gap is what the professionals know that this model doesn't. StatsBomb says its model adds the keeper's position and status, every attacker and defender in the frame, and the height at which the ball was struck, on top of the usual distance, angle, body part and type of assist. Ours has nine features and no assists.
Why it matters
- The right data beats a cleverer method. The same logistic regression went from 0.294 to 0.266 just by being shown where people stood. No new algorithm was needed.
- A model is a measuring tool. Its weights put numbers on things fans argue about: a defender in the line, a header instead of a volley.
- Watch what it can't separate. The keeper result is a pattern in the data, not advice to keepers to stay on their line.
- Expected goals add up. Give every shot in a match its xG and Did we deserve to win? turns them into each side's chances of winning, just as the Bernoulli distribution showed.
Limitations
- Only the players near the ball. The freeze frame records those around the shot, so a defender out of shot may simply be missing.
- No shot quality. It doesn't know how hard or where the ball was hit, the shooter's body shape, or whether the ball was bouncing.
- Every shooter is average. A top striker scores more than this from the same spot; the model can't tell who's shooting.
- Open play only, and not Scotland. No penalties, free kicks or corners, and the matches are top leagues and tournaments, not the Scottish Premiership.
- One random split. Cross-validation over several splits would be steadier.
Try it yourself
Open the Shot Lab. Start from Edge of the box, then drag the three defenders out of the triangle and watch the xG climb from 0.07. Switch the model to Distance only and it can't see the difference at all. Try One on one and drag the keeper back to the line: the number falls, which is the keeper puzzle above, live. Or press Watch a move and see four short attacks play out, the numbers rising and falling as the players move.
Or by hand: next time you watch a shot from the edge of the box, count the defenders between the ball and the posts before it's struck. One or more, and the model says the chance just got smaller.
Reproduce the analysis
You'll need StatsBomb's open data on your computer, about 15 GB unpacked (the download is compressed). This fetches only the events and the match lists:
Show the Bash3 lines, ready to copy and run.
mkdir -p ~/statsbomb-open-data && curl -sL https://codeload.github.com/statsbomb/open-data/tar.gz/refs/heads/master \
| tar xz -C ~/statsbomb-open-data --strip-components=1 --wildcards 'open-data-master/data/competitions.json' \
'open-data-master/data/matches/*' 'open-data-master/data/events/*'
Then run this. It needs nothing beyond standard Python and takes about ten minutes, most of it reading the files and fitting:
Show the Python165 lines, ready to copy and run.
# From Football Data Science by Bryan McGuire. Free to use with credit.
# https://www.footballdatascience.co.uk/learn/expected-goals-from-scratch
import json
import os
import random
from math import asin, atan2, exp, hypot, log, sqrt
DATA = os.path.expanduser("~/statsbomb-open-data/data")
YD = 0.9144 # StatsBomb's pitch is 120 x 80 yards; distances are reported in metres
def load(path):
with open(os.path.join(DATA, path), encoding="utf-8") as f:
return json.load(f)
# every open-play shot with a freeze frame that shows the keeper
shots = []
for c in load("competitions.json"):
for m in load(f"matches/{c['competition_id']}/{c['season_id']}.json"):
for e in load(f"events/{m['match_id']}.json"):
s = e.get("shot")
if not s or s["type"]["name"] != "Open Play" or "freeze_frame" not in s:
continue
opponents = [p for p in s["freeze_frame"] if not p["teammate"]]
keeper = [p["location"] for p in opponents if p["position"]["name"] == "Goalkeeper"]
if keeper:
shots.append({"match": m["match_id"], "women": c["competition_gender"] == "female",
"ball": e["location"][:2], "keeper": keeper[0][:2], "head": s["body_part"]["name"] == "Head",
"defenders": [p["location"][:2] for p in opponents if p["position"]["name"] != "Goalkeeper"],
"goal": s["outcome"]["name"] == "Goal", "statsbomb": s["statsbomb_xg"]})
def inside(p, a, b, c):
"""Is point p inside the triangle a, b, c?"""
side = lambda p, q, r: (p[0] - r[0]) * (q[1] - r[1]) - (q[0] - r[0]) * (p[1] - r[1])
d = side(p, a, b), side(p, b, c), side(p, c, a)
return not (min(d) < 0 < max(d))
def visible(ball, keeper, defenders, r=0.5):
"""Radians of the goal mouth the shooter can see past the bodies, each blocking r yards either side."""
x, y = ball
a1, a2 = sorted((atan2(36 - y, 120 - x), atan2(44 - y, 120 - x)))
blocks = []
for px, py in [keeper] + defenders:
d = hypot(px - x, py - y)
if px <= x: # behind the ball: blocks nothing
continue
if d <= r: # touching the ball: blocks everything
blocks.append((a1, a2))
continue
c, w = atan2(py - y, px - x), asin(r / d)
blocks.append((c - w, c + w))
seen, cur = 0.0, a1
for lo, hi in sorted(blocks): # walk across the goal, adding up the gaps between bodies
if hi <= cur:
continue
if lo > cur:
seen += min(lo, a2) - cur
cur = max(cur, hi)
if cur >= a2:
break
return seen + max(0.0, a2 - cur)
def features(ball, keeper, defenders, head):
"""What the model knows about a shot. The goal is at x = 120, its posts at y = 36 and 44."""
x, y = ball
gx, gy = 120 - x, 40 - y
g = hypot(gx, gy)
in_way = lambda p: inside(p, ball, (120, 36), (120, 44)) # inside the triangle from the ball to the posts
return {"dist": g * YD, # metres to the centre of the goal
"angle": abs(atan2(36 - y, gx) - atan2(44 - y, gx)), # how much of the goal it can see
"head": int(head),
"blockers": sum(in_way(p) for p in defenders), # defenders in the way
"near": min([hypot(p[0] - x, p[1] - y) * YD for p in defenders] + [30]), # metres to the nearest one
"gk_in": int(in_way(keeper)), # keeper in the way
"gk_off": abs(gx * (keeper[1] - y) - gy * (keeper[0] - x)) / g * YD, # keeper off the ball-goal line
"gk_line": (120 - keeper[0]) * YD, # keeper off the goal line
"visible": visible(ball, keeper, defenders)} # goal you can see past them
for s in shots:
s["f"] = {**features(s["ball"], s["keeper"], s["defenders"], s["head"]), "women": int(s["women"])}
# three models, each knowing more; a term is a feature, sometimes logged or capped
DISTANCE = [{"f": "dist"}]
ANGLE = DISTANCE + [{"f": "angle"}]
EVERYTHING = [{"f": "dist", "log": True}, {"f": "dist"}, {"f": "angle"}, {"f": "head"}, {"f": "blockers"},
{"f": "near", "cap": 3}, {"f": "gk_in"}, {"f": "gk_off"}, {"f": "gk_line"}, {"f": "visible"}]
def value(f, t):
v = f[t["f"]]
return log(v) if t.get("log") else min(v, t["cap"]) if "cap" in t else v
def fit(terms, data, steps=1000, rate=1.0):
"""Logistic regression by gradient descent on log loss, each feature measured in standard deviations."""
X = [[value(s["f"], t) for t in terms] for s in data]
k = len(terms)
mean = [sum(x[j] for x in X) / len(X) for j in range(k)]
sd = [sqrt(sum((x[j] - mean[j]) ** 2 for x in X) / len(X)) for j in range(k)]
Z = [[(x[j] - mean[j]) / sd[j] for j in range(k)] for x in X]
w = [0.0] * (k + 1)
for _ in range(steps):
grad = [0.0] * (k + 1)
for z, s in zip(Z, data):
miss = 1 / (1 + exp(-(w[0] + sum(a * b for a, b in zip(w[1:], z))))) - s["goal"]
grad[0] += miss
for j in range(k):
grad[j + 1] += miss * z[j]
w = [a - rate * g / len(Z) for a, g in zip(w, grad)]
return {"intercept": w[0], "terms": [{**t, "w": w[j + 1], "mean": mean[j], "sd": sd[j]} for j, t in enumerate(terms)]}
def predict(model, f):
z = model["intercept"] + sum(t["w"] * (value(f, t) - t["mean"]) / t["sd"] for t in model["terms"])
return 1 / (1 + exp(-z))
def scores(ps, data):
loss = [-log(p if s["goal"] else 1 - p) for p, s in zip(ps, data)]
mean = sum(loss) / len(loss)
se = sqrt(sum((l - mean) ** 2 for l in loss) / len(loss) / len(loss))
return mean, se, sum((p - s["goal"]) ** 2 for p, s in zip(ps, data)) / len(data)
# split by match: 60% to train, 20% held back for choices, 20% kept for one final test
matches = sorted({s["match"] for s in shots})
random.Random(2026).shuffle(matches)
part = {m: "train" if i < 0.6 * len(matches) else "held" if i < 0.8 * len(matches) else "test" for i, m in enumerate(matches)}
train, held, test = ([s for s in shots if part[s["match"]] == p] for p in ("train", "held", "test"))
goal_rate = sum(s["goal"] for s in train + held) / len(train + held)
print(f"{len(shots)} shots from {len(matches)} matches: train {len(train)}, held back {len(held)}, test {len(test)}")
print(f"{goal_rate:.1%} went in; {sum(s['women'] for s in shots) / len(shots):.0%} of the shots are from women's matches")
# two questions settled on the held-back matches: do women's shots need their own term,
# and does "keeper off the line" earn its place?
for name, terms in (("everything", EVERYTHING), ("plus a women's-match term", EVERYTHING + [{"f": "women"}]),
("without keeper off the line", [t for t in EVERYTHING if t["f"] != "gk_line"])):
m = fit(terms, train)
print(f"held back, {name}: log loss {scores([predict(m, s['f']) for s in held], held)[0]:.4f}")
# refit on train + held back, then the one look at the test matches
final = {name: fit(terms, train + held) for name, terms in (("distance only", DISTANCE), ("distance and angle", ANGLE),
("everything", EVERYTHING))}
for name, ps in [("guess the goal rate", [goal_rate] * len(test))] + [(n, [predict(m, s["f"]) for s in test]) for n, m in final.items()] \
+ [("StatsBomb's xG", [s["statsbomb"] for s in test])]:
loss, se, brier = scores(ps, test)
print(f"test, {name:20}: log loss {loss:.3f} (± {2 * se:.3f}), Brier {brier:.3f}")
best = final["everything"]
print("weights, per standard deviation:", ", ".join(f"{t['f']}{' (log)' if t.get('log') else ''} {t['w']:+.2f}" for t in best["terms"]))
# made-up situations, as in the Shot Lab (yards; goal at x = 120)
SITUATIONS = {
"edge of the box, three defenders": ((101, 41), (118.5, 40), [(105, 40), (104, 45), (107, 35)], False),
"same, all three out of the way": ((101, 41), (118.5, 40), [(105, 47), (104, 49), (107, 31)], False),
"one on one, keeper rushing out": ((108, 38), (116, 39.5), [(103, 41)], False),
"one on one, keeper on the line": ((108, 38), (119.5, 39.5), [(103, 41)], False),
"six yards out, foot": ((114, 40), (119, 40), [], False),
"six yards out, header": ((114, 40), (119, 40), [], True),
}
for name, (ball, keeper, defenders, head) in SITUATIONS.items():
print(f"{name:34}: " + ", ".join(f"{n} {predict(m, features(ball, keeper, defenders, head)):.2f}" for n, m in final.items()))
Further reading
- StatsBomb open data, on GitHub (now under Hudl, which owns StatsBomb): the free event data used here, with documentation of every field, including the shot freeze frame.
- What are Expected Goals (xG)?, Hudl StatsBomb's explainer: what usually goes into an xG model, and what their own model adds.
- What Is Expected Goals (xG)?, Opta Analyst: another provider's explanation, with examples of how xG is used in match analysis.
- Expected goals, Wikipedia: a short history of the idea and its spread into broadcasting.