How many shots is a goal worth? Linear regression
Draw the best straight line through 310 Scottish team-seasons and it says a goal comes every 3.1 shots on target. Then test it on seasons it never saw, and watch one version break when the way shots are counted changes.
Beginner Part 16 of Machine Learning Through Football
New to the notation? The symbols explained
Contents
The football question
Shots get called a vanity stat, but over a season the shooters are the scorers. So how many shots is a goal actually worth? And if a team takes two more shots a game, how many more goals should it expect?
Answering that with one number is the job of linear regression: the oldest model in data science, and still one of the most used.
The concept
Plot every team-season as a point: shots on target a game along the bottom, goals a game up the side. Then draw the straight line that runs closest to all of them:
$$\text{goals} = a + b \times \text{shots on target}$$
In plain football
- b, the slope, is what one more shot on target a game is worth in goals a game. It's the answer to the question.
- a, the start, is where the line would cross zero shots. It rarely means anything on its own; it just puts the line in the right place.
- Both are worked out from the data. Nobody picks them by hand.
"Closest to all of them" needs a definition. For each team-season, the miss is the gap between its real goals and the line's. Linear regression squares every miss and picks the line that makes their total as small as possible, which is why it's also called least squares. Squaring stops misses above and below the line cancelling out, and makes one big miss count more than several small ones.
The same idea gave how many points a goal is worth, and rolling downhill found a one-dial line by trial and error. For a straight line there's also a direct formula, and that's what the code below uses.
The data: every Scottish Premiership team-season from 2000/01 to 2025/26 with shots recorded in at least 30 matches, 310 team-seasons, from football-data.co.uk. The average side scored 1.34 goals a game, about 51 over a 38-game season.
One line, two ways
Fit one line using all shots, and another using only shots on target:
| Line | One goal every | Explains |
|---|---|---|
| Goals from shots | 7.1 shots | 61% |
| Goals from shots on target | 3.1 on target | 72% |
Explains is the share of the differences in goals between team-seasons that the line accounts for, often written R². The rest is everything the line doesn't see: finishing, the quality of the chances, luck. The two figures are the shots myth's correlations, 0.78 and 0.85, squared: for a single straight line, that's always how they're linked.
Shots on target win. Where the shots go matters as well as how many there are.
Two features at once
Linear regression can use more than one feature. Split every team's shots into those on target and those off target, blocked or wide, and fit both at once:
$$\begin{aligned} \text{goals} = \;&{-0.360} \\ &+ 0.290 \times \text{on target} \\ &+ 0.063 \times \text{off target} \end{aligned}$$
In plain football
- Each shot on target a game is worth 0.29 goals a game, with the shots off target held the same.
- Each shot off target is still worth 0.06, about one goal in 16. A miss isn't worthless: a team that gets enough shots away, even off target, is usually a team that gets forward.
- Together they explain 78% of the differences, more than either alone.
Each weight now means "with the other feature held the same". That's the big step from one feature to several, and it's how logistic regression weighs its three gaps.
The honest test
So far the lines were fitted to all 310 team-seasons and judged on the same ones. As in part 2, the real question is how they do on seasons they never saw. Fit on the 250 team-seasons from 2000/01 to 2020/21, then predict the 60 from 2021/22 to 2025/26:
| Line | Typical miss, a season | Explains |
|---|---|---|
| Guess the league average | 20 goals | – |
| Shots | 16.4 goals | 30% |
| Shots on target | 10.5 goals | 72% |
| On and off target | 7.4 goals | 86% |
The typical miss is the square root of the average squared miss, turned into goals over 38 games.
Two of the lines hold up: shots on target explained 71% of the training seasons and 72% of the test seasons. The shots line doesn't: 65% of the training seasons, only 30% of the test seasons. The world changed under it. Here's what the average side recorded each season:
From 2020/21, the average side was recorded with about two more shots a game: 9.8 in 2019/20, 11.7 the next season and around 12.6 since. Goals didn't follow; 2021/22 had 1.23 a game, lower than most seasons before. A line that learned "about 7 shots a goal" from the old seasons expects far too many goals from the new ones. Earlier, from 2013/14, shots on target dropped from around 5 a game to around 4.2, while each one became more likely to go in, about 0.32 goals against about 0.27.
Steps that sudden look more like changes in how shots were recorded than changes in the football, though the data can't say which. Either way, a line is only as good as the world staying the same, and only a test on unseen seasons catches it when it doesn't. The two-feature line does best on the test, but be careful: part of that is luck in how the two changes happen to offset each other.
Above and below the line
The miss for each team-season, turned into goals over the season, is how many more or fewer it scored than its shots on target suggest:
| Team-season | Against the line |
|---|---|
| Celtic 2022/23 | +36 goals |
| Celtic 2024/25 | +30 goals |
| Celtic 2013/14 | +26 goals |
| Dunfermline 2004/05 | −21 goals |
| Falkirk 2007/08 | −22 goals |
| Celtic 2009/10 | −22 goals |
Does beating the line carry over to the next season? At first sight yes: a correlation of 0.42. But much of that is the recording change. On average, team-seasons from 2013/14 on sit 4.9 goals a season above the line and those before it 4.9 below, and next season is nearly always in the same era as this one. Take out what every team did that season, and it drops to 0.24 (luck margin ±0.12). Without Celtic and Rangers it's 0.15 (±0.13), barely clear of luck.
So a little of it is real, and the clearest part is Celtic and Rangers: against their own season's average they beat the line by 6.2 goals a season. Shots on target don't measure how good the chances were, and dominant sides get closer, clearer chances. For everyone else, beating the line is mostly a good year, much as beating your goal difference is.
Why it matters
- One number with a meaning. "A shot on target is worth about 0.3 goals" is something a coach, a scout or a fan can use. Most models can't be read like that.
- Several features, each held fair. The off-target weight only makes sense with on target held the same. That idea runs through almost every model in this series.
- Test on unseen seasons, every time. The shots line looked fine on the data it was fitted to. Only the test showed it had gone out of date.
- Check what else changes. "Beating the line carries over" looked like a skill until the change in recording was taken out. A pattern can come from something that moves along with what you measure.
Limitations
- Shots, not chances. A tap-in and a 30-yard hopeful count the same. Expected goals (xG) would weigh each chance, but this data doesn't have it.
- Straight lines only. The bands sit close to the line, but a real effect could bend at the extremes, where there are few team-seasons.
- How the data was recorded. The two steps in the shot counts change what a "shot" means between eras. A line fitted on one era can mislead in another.
- One league, season averages. Over a handful of matches luck swamps everything here.
Try it yourself
The Goals from Shots Calculator uses the two-feature line fitted on the training seasons: 0.283 goals for each shot on target and 0.073 for each off target, a touch different from the all-seasons line above because it saw fewer, older seasons. Enter your team's shots and see how many goals they should bring over a season, and how often a match brings none, one, or two or more.
Or by hand: take your team's shots on target per game from last season, multiply by 0.32, take off 0.13, and multiply by 38. How close is it to the goals they actually scored? Above or below the line?
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It runs in under a second:
Show the Python103 lines, ready to copy and run.
import csv
from collections import defaultdict
from math import sqrt
from statistics import correlation, mean
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
COLUMNS = ("FTHG", "FTAG", "HS", "AS", "HST", "AST")
# every team-season with shots recorded in at least 30 matches: goals, shots and shots on target per game
totals = defaultdict(lambda: defaultdict(int))
for s in names:
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
for r in csv.DictReader(f):
if not all(r.get(c) for c in COLUMNS):
continue
for team, goals, shots, on_target in ((r["HomeTeam"], r["FTHG"], r["HS"], r["HST"]),
(r["AwayTeam"], r["FTAG"], r["AS"], r["AST"])):
t = totals[(s, team)]
t["n"] += 1
t["goals"] += int(goals)
t["shots"] += int(shots)
t["on"] += int(on_target)
rows = [{"season": s, "team": team, "n": t["n"], "goals": t["goals"] / t["n"], "shots": t["shots"] / t["n"],
"on": t["on"] / t["n"], "off": (t["shots"] - t["on"]) / t["n"]}
for (s, team), t in totals.items() if t["n"] >= 30]
print(f"{len(rows)} team-seasons, {len({r['season'] for r in rows})} seasons; the average side: "
f"{mean(r['goals'] for r in rows):.2f} goals, {mean(r['shots'] for r in rows):.1f} shots, {mean(r['on'] for r in rows):.1f} on target a game")
def fit(data, features): # least squares: the start and weights that make the squared misses smallest
X = [[1] + [r[f] for f in features] for r in data]
y = [r["goals"] for r in data]
k = len(X[0])
# the normal equations, (X'X) w = X'y, solved by elimination
M = [[sum(x[i] * x[j] for x in X) for j in range(k)] + [sum(x[i] * v for x, v in zip(X, y))] for i in range(k)]
for i in range(k):
for j in range(k):
if j != i:
ratio = M[j][i] / M[i][i]
M[j] = [a - ratio * b for a, b in zip(M[j], M[i])]
return [M[i][k] / M[i][i] for i in range(k)]
def predict(w, r, features):
return w[0] + sum(wj * r[f] for wj, f in zip(w[1:], features))
def explained(w, data, features): # R squared: the share of the spread in goals the line accounts for
y = [r["goals"] for r in data]
miss = sum((r["goals"] - predict(w, r, features)) ** 2 for r in data)
return 1 - miss / sum((v - mean(y)) ** 2 for v in y)
def typical_miss(w, data, features): # root mean squared miss, in goals a game
return sqrt(mean([(r["goals"] - predict(w, r, features)) ** 2 for r in data]))
MODELS = {"shots": ["shots"], "on target": ["on"], "on and off target": ["on", "off"]}
# 1. the lines through all 310 team-seasons
for name, features in MODELS.items():
w = fit(rows, features)
weights = ", ".join(f"{f} {wj:+.3f}" for f, wj in zip(features, w[1:]))
print(f"goals from {name}: start {w[0]:+.3f}, {weights}; explains {explained(w, rows, features):.0%}")
print(f"correlations with goals: shots {correlation([r['shots'] for r in rows], [r['goals'] for r in rows]):.2f}, "
f"on target {correlation([r['on'] for r in rows], [r['goals'] for r in rows]):.2f}")
w = fit(rows, ["on"])
print(f"one goal every {1 / fit(rows, ['shots'])[1]:.1f} shots, every {1 / w[1]:.1f} shots on target")
for lo in (2.5, 3.5, 4.5, 5.5, 6.5, 7.5): # average goals for team-seasons grouped by shots on target, for the chart
band = [r for r in rows if lo <= r["on"] < lo + 1]
print(f" on target {lo}-{lo + 1}: {len(band)} team-seasons, goals {mean(r['goals'] for r in band):.2f}, line {w[0] + w[1] * (lo + 0.5):.2f}")
# 2. the honest test: fit on 2000/01-2020/21, predict the five seasons since
train = [r for r in rows if r["season"] < "2122"]
test = [r for r in rows if r["season"] >= "2122"]
average = sqrt(mean([(r["goals"] - mean(t["goals"] for t in train)) ** 2 for r in test]))
print(f"train {len(train)}, test {len(test)}; guess the average: misses by {average:.3f} a game, {38 * average:.0f} a season")
for name, features in MODELS.items():
w = fit(train, features)
miss = typical_miss(w, test, features)
print(f" {name}: misses by {miss:.3f} a game, {38 * miss:.1f} a season; explains {explained(w, train, features):.0%} "
f"of the training seasons, {explained(w, test, features):.0%} of the test seasons")
print("calculator line:", ", ".join(f"{v:+.4f}" for v in fit(train, ["on", "off"])))
# 3. what each season recorded
for s in names:
season = [r for r in rows if r["season"] == s]
print(f" 20{s[:2]}/{s[2:]}: shots {mean(r['shots'] for r in season):.1f}, on target {mean(r['on'] for r in season):.1f}, "
f"goals {mean(r['goals'] for r in season):.2f}, goals per shot on target {sum(r['goals'] for r in season) / sum(r['on'] for r in season):.2f}")
# 4. above and below the line, in goals over the season, and whether it carries over
w = fit(rows, ["on"])
beat = {(r["season"], r["team"]): (r["goals"] - predict(w, r, ["on"])) for r in rows}
seasonal = sorted(((v * r["n"], r["season"], r["team"]) for r, v in zip(rows, beat.values())))
for v, s, team in seasonal[:3] + seasonal[-3:]:
print(f" {team} 20{s[:2]}/{s[2:]}: {v:+.0f} goals against the line")
for era, inside in (("before 2013/14", lambda s: s < "1314"), ("2013/14 on", lambda s: s >= "1314")):
print(f"average against the line, {era}: {38 * mean(v for (s, _), v in beat.items() if inside(s)):+.1f} goals a season")
season_avg = {s: mean(v for (ss, _), v in beat.items() if ss == s) for s in names}
fair = {k: v - season_avg[k[0]] for k, v in beat.items()} # against the line, less what every team did that season
old_firm = [v * 38 for (s, team), v in fair.items() if team in ("Celtic", "Rangers")]
print(f"Celtic and Rangers, against their season's average: {mean(old_firm):+.1f} goals a season ({len(old_firm)} team-seasons)")
for label, d in (("raw", beat), ("season average removed", fair)):
for who, keep in (("all teams", lambda t: True), ("without Celtic and Rangers", lambda t: t not in ("Celtic", "Rangers"))):
pairs = [(v, d[(names[names.index(s) + 1], team)]) for (s, team), v in d.items()
if keep(team) and s != names[-1] and (names[names.index(s) + 1], team) in d]
print(f"carries over, {label}, {who}: correlation {correlation(*zip(*pairs)):.2f} "
f"(luck margin +/- {2 / sqrt(len(pairs) - 3):.2f}, {len(pairs)} pairs)")
Further reading
- An Introduction to Statistical Learning, by James, Witten, Hastie and Tibshirani: a free textbook whose chapter on linear regression covers least squares, R² and several features at once, with worked examples.
- Regression analysis, Seeing Theory, Brown University: an interactive page where you drag points and watch the least-squares line and its squared misses change.
- Linear models, scikit-learn user guide: the standard Python library's documentation, starting with ordinary least squares.