Matches like this one. K-nearest neighbours
To forecast a match, find the past matches most like it and see what happened. No training, no weights, just the right number of lookalikes. On five SPFL test seasons, the 300 nearest draw level with logistic regression.
Intermediate Part 15 of Machine Learning Through Football
New to the notation? The symbols explained
Contents
The football question
Before a big match, someone always says "the last time these two met in this kind of form…" and reaches for a memory. What if, instead of one memory, you looked up every past match that was set up the same way, and counted what happened?
That's k-nearest neighbours, or KNN: describe the match, find the k past matches most like it, and let them vote.
The concept
Every model in this series so far has learned something: logistic regression learned weights, a decision tree learned questions. KNN learns nothing. The past matches are the model. All it needs is a way to say how alike two matches are, and that's the distance from measuring player similarity, with matches in place of players.
Each match is described by the same three numbers as parts 8 to 11, each the home side's figure minus the away side's, in points a game: recent form, this season so far and last season. Two matches are close if all three gaps are close:
$$d = \sqrt{\Delta_1^2 + \Delta_2^2 + \Delta_3^2}$$
In plain football
- Δ₁ is how far apart the two matches' form gaps are, Δ₂ their this-season gaps, Δ₃ their last-season gaps.
- Each gap is first measured in standard deviations, so that one feature with bigger numbers can't drown out the others. That's the scale trap again.
- d = 0 means two matches with exactly the same gaps. The bigger d, the less alike they are.
To forecast a match, KNN measures its distance to every one of the 3,900 training matches, 2001/02 to 2020/21, takes the k nearest, and counts:
$$P(\text{home}) = \frac{H + h}{k + 1}$$
In plain football
- k is how many neighbours vote: the one setting KNN has.
- H is how many of those k neighbours were home wins.
- h is the share of home wins across all training matches, about 0.44. It adds one imaginary match, split between home win, draw and away win at their usual rates, so that no result is ever given a chance of zero. A zero would make the log loss infinite the first time that result happened.
- Draws and away wins are counted the same way, and the three chances add up to 1.
One real match
Hearts 1–0 Hibernian, 4 October 2025. Before kick-off, Hearts had taken 1.60 points a game more than Hibs over their last five, and 1.33 a game more over the season so far, though Hibs had been 0.16 a game better the season before. Its five nearest training matches:
| Season | Match | Gaps |
|---|---|---|
| 2007/08 | Motherwell 3–0 Gretna | +1.80, +1.29, −0.17 |
| 2012/13 | Hibernian 2–2 Inverness | +1.60, +1.17, −0.16 |
| 2013/14 | Aberdeen 1–0 Ross County | +1.60, +1.13, −0.13 |
| 2002/03 | Kilmarnock 2–2 Aberdeen | +1.80, +1.17, −0.16 |
| 2014/15 | Partick 1–2 St Mirren | +1.40, +1.40, −0.03 |
None of them is an Edinburgh derby, and that's the point: KNN never sees team names, only how far apart the two sides have been. With the 300 nearest voting, it gave Hearts 57.3%, a draw 21.3% and Hibs 21.4%. Logistic regression said 62.0%, 20.1% and 17.9%.
How many neighbours?
The whole method rests on k, so it has to be chosen honestly: on the held-back seasons, 2016/17 to 2020/21, with the neighbours drawn from 2001/02 to 2015/16. The test seasons stay untouched.
| Neighbours | Held-back log loss |
|---|---|
| 5 | 1.158 |
| 20 | 1.008 |
| 100 | 0.976 |
| 300 | 0.9705 |
| 800 | 0.984 |
| Base rates | 1.073 |
Five neighbours do worse than knowing nothing. Five matches are mostly luck: four home wins among them makes the forecast (4 + 0.44) ÷ 6, a 74% chance of a home win, on evidence you could fit on a bench. That's too jumpy, high variance, and the same overfitting a deep decision tree showed in part 9.
Too many neighbours is too stubborn. At 800, the neighbours reach further and further from the match being forecast, so every forecast drifts back towards the league average. That's high bias.
In between, 300 is best. It sounds like a lot, but it's about a tenth of the matches the neighbours were drawn from, and for an evenly matched game all 300 sit within half a standard deviation of it. The curve is fairly flat from 100 to 400, so the exact number matters less than avoiding the extremes.
Against the rest
Now the one look at the 990 test matches, 2021/22 to 2025/26, with the neighbours drawn from all 3,900 training matches:
| Model | Test accuracy | Test log loss |
|---|---|---|
| Base rates | 47.2% | 1.056 |
| Tree, two questions | 55.5% | 0.954 |
| Forest, groups of 100+ | 54.3% | 0.953 |
| Gradient boosting | 54.5% | 0.951 |
| Logistic regression | 54.8% | 0.950 |
| KNN, 300 neighbours | 55.4% | 0.949 |
| Dixon-Coles | 55.3% | 0.945 |
| Bookmaker | 56.2% | 0.932 |
KNN draws level with logistic regression. It's 0.0013 ahead on the test seasons, but the luck margin on that gap is ±0.0065, five times bigger, and on the held-back seasons logistic regression was ahead instead (0.968 against 0.9705). Call it a draw. Like every model before it, KNN never makes a draw the favourite: 663 home wins and 327 away wins picked, no draws.
So the simplest model in the series, one that learns nothing at all, matches the best of the four that do. It's more evidence for what gradient boosting found: on these three numbers, the features are the ceiling. Getting past it took a different kind of model, Dixon-Coles, which rates every team's attack and defence.
An evenly matched game
Where KNN does stand out is the evenly matched game, every gap 0:
| Model | Home | Draw | Away |
|---|---|---|---|
| Logistic regression | 43.5% | 25.8% | 30.7% |
| KNN, 300 neighbours | 37.7% | 29.3% | 33.0% |
| Close matches, training seasons | 38.5% | 28.5% | 33.0% |
The last row, from part 10, is what actually happened in the 397 training matches where the sides were within 0.2 points a game of each other on both season gaps. KNN lands almost on it, and that's no coincidence: its forecast is what happened in nearby matches. Logistic regression has to fit one straight line through every match, so it pushes the home side too high in the middle. KNN only ever looks locally.
Why it matters
- Similarity is a model. "Find me matches like this one" is how recruitment shortlists replacements, how recommendation engines suggest the next thing, and how analysts pick historical comparisons. KNN is that idea, scored honestly.
- One setting, chosen honestly, is the whole job. KNN has no weights to learn, only k. Pick it on the test seasons and you'd be fooling yourself; pick it on held-back seasons and the test score means something.
- Local beats global where the pattern bends. A straight-line model gets the middle wrong; a model that only asks the neighbours doesn't have to commit to a line.
- Simple can match clever. A method that learns nothing drew level with every learner in the series. Try the simple thing first.
Limitations
- All features count the same. Distance gives recent form as much say as the two season gaps, although logistic regression found it adds almost nothing. Here, dropping it made almost no difference (0.9715 without it against 0.9705 with it, on the held-back seasons), but with a useless feature among many, KNN can't learn to ignore it.
- Slow to forecast. Every forecast measures the distance to all 3,900 past matches. Fine here, slow for millions.
- Thin at the edges. For an evenly matched game the 300th neighbour is 0.48 standard deviations away. For a big mismatch, gaps of +2.4, +2.0 and +1.8, it's 2.32, nearly five times as far, so the "nearest" 300 include some quite different games.
- Three features, one league. No injuries, team news or managers, and only the Scottish Premiership, where two clubs have usually been far ahead of the rest.
- One held-back split. k was chosen on one block of five seasons. Cross-validation over several blocks would be steadier.
Try it yourself
The Similar Matches Finder does this live: set the three gaps and it lists the 20 nearest past matches and what happened in them. Set k to 5 and watch the forecast lurch as you nudge a slider; set it back to 300 and it settles.
Or by hand: pick this weekend's most even-looking fixture, then find five past matches between sides that were level on points at the same stage of the season. How many were home wins? Now find twenty. Did the answer settle down?
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It takes about half a minute, most of it measuring distances:
Show the Python127 lines, ready to copy and run.
import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import exp, log, sqrt
POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
RESULTS = "HDA"
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
def season(s):
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
return games
def points_per_game(games):
pts, n = Counter(), Counter()
for r in games:
for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
pts[team] += p
n[team] += 1
return {t: pts[t] / n[t] for t in n}
# the same three features as part 8, each the home side's figure minus the away side's, all known before kick-off
rows = []
for s_last, s in zip(names, names[1:]):
last = points_per_game(season(s_last))
promoted = 0.85 * sum(last.values()) / len(last)
history = defaultdict(list)
for r in season(s):
h, a = r["HomeTeam"], r["AwayTeam"]
if len(history[h]) >= 5 and len(history[a]) >= 5:
rows.append({"season": s, "match": f"{h} {r['FTHG']}-{r['FTAG']} {a}", "date": r["Date"], "y": r["FTR"], "x": [
(sum(history[h][-5:]) - sum(history[a][-5:])) / 5, # points a game, last five
sum(history[h]) / len(history[h]) - sum(history[a]) / len(history[a]), # points a game this season
last.get(h, promoted) - last.get(a, promoted)]}) # points a game last season
for team, p in zip((h, a), POINTS[r["FTR"]]):
history[team].append(p)
def standardise(fit, other): # in standard deviations of the fitting matches, so each feature counts the same
mean = [sum(r["x"][j] for r in fit) / len(fit) for j in range(3)]
sd = [sqrt(sum((r["x"][j] - mean[j]) ** 2 for r in fit) / len(fit)) for j in range(3)]
for r in fit + other:
r["z"] = [(r["x"][j] - mean[j]) / sd[j] for j in range(3)]
return mean, sd
def distance(a, b):
return sqrt(sum((p - q) ** 2 for p, q in zip(a, b)))
def knn(fit, z, k, use=slice(0, 3)): # the share of each result among the k nearest, plus one imaginary match at the base rates
base = Counter(r["y"] for r in fit)
near = sorted(fit, key=lambda r: distance(r["z"][use], z[use]))[:k]
seen = Counter(r["y"] for r in near)
return {c: (seen[c] + base[c] / len(fit)) / (k + 1) for c in RESULTS}, near
def logistic(fit): # part 8's model, refitted here so the two can be compared match by match
def predict(w, z):
e = {c: exp(w[c][0] + sum(wj * zj for wj, zj in zip(w[c][1:], z))) for c in RESULTS}
return {c: e[c] / sum(e.values()) for c in RESULTS}
w = {c: [0.0] * 4 for c in RESULTS}
for _ in range(200):
grad = {c: [0.0] * 4 for c in "HA"}
for r in fit:
p = predict(w, r["z"])
for c in "HA":
miss = p[c] - (c == r["y"])
grad[c] = [g + miss * v for g, v in zip(grad[c], [1] + r["z"])]
for c in "HA":
w[c] = [wj - gj / len(fit) for wj, gj in zip(w[c], grad[c])]
return lambda z: predict(w, z)
def log_loss(forecasts, data):
return sum(-log(p[r["y"]]) for p, r in zip(forecasts, data)) / len(data)
def base_rates(fit, data):
n = Counter(r["y"] for r in fit)
return log_loss([{c: n[c] / len(fit) for c in RESULTS}] * len(data), data)
# 1. choose k on the held-back seasons 2016/17-2020/21, fitting on 2001/02-2015/16
fit = [r for r in rows if r["season"] < "1617"]
held = [r for r in rows if "1617" <= r["season"] < "2122"]
standardise(fit, held)
print(f"fitting {len(fit)} matches, held back {len(held)}; base rates {base_rates(fit, held):.4f}; "
f"home wins {sum(r['y'] == 'H' for r in fit) / len(fit):.3f} of the fitting matches")
scores = {}
for k in (5, 10, 20, 50, 100, 200, 300, 400, 600, 800):
scores[k] = log_loss([knn(fit, r["z"], k)[0] for r in held], held)
print(f"k = {k:3}: held-back log loss {scores[k]:.4f}")
best = min(scores, key=scores.get)
no_form = log_loss([knn(fit, r["z"], best, use=slice(1, 3))[0] for r in held], held) # distance on the two season gaps only
print(f"k = {best} without recent form: held-back log loss {no_form:.4f}")
lr_held = logistic(fit)
print(f"chosen: k = {best}; logistic regression on the same seasons {log_loss([lr_held(r['z']) for r in held], held):.4f}")
# 2. refit on all 3,900 training matches and score the 990 test matches once
train = [r for r in rows if r["season"] < "2122"]
test = [r for r in rows if r["season"] >= "2122"]
mean, sd = standardise(train, test)
lr = logistic(train)
p_knn = [knn(train, r["z"], best)[0] for r in test]
p_lr = [lr(r["z"]) for r in test]
picks = Counter(max(RESULTS, key=p.get) for p in p_knn)
right = sum(max(RESULTS, key=p.get) == r["y"] for p, r in zip(p_knn, test))
print(f"test, {len(test)} matches: base rates {base_rates(train, test):.3f}, logistic regression {log_loss(p_lr, test):.4f}, "
f"KNN {log_loss(p_knn, test):.4f}, KNN accuracy {right / len(test):.1%}, picks {dict(picks)}")
gap = [log(q[r["y"]]) - log(p[r["y"]]) for p, q, r in zip(p_knn, p_lr, test)] # KNN's log loss minus LR's, match by match
avg = sum(gap) / len(gap)
se = sqrt(sum((g - avg) ** 2 for g in gap) / (len(gap) - 1) / len(gap))
print(f"KNN minus logistic regression: {avg:+.4f}, luck margin +/- {2 * se:.4f}")
# 3. an evenly matched game: every gap 0
even = [(0 - mean[j]) / sd[j] for j in range(3)]
p, near = knn(train, even, best)
q = lr(even)
print("even game, KNN:", ", ".join(f"{c} {p[c]:.1%}" for c in RESULTS), "| logistic:", ", ".join(f"{c} {q[c]:.1%}" for c in RESULTS),
f"| furthest of the {best}: {distance(near[-1]['z'], even):.2f} sd")
mismatch = [(g - mean[j]) / sd[j] for j, g in enumerate((2.4, 2.0, 1.8))] # a big mismatch, home side far better
print(f"big mismatch (+2.4, +2.0, +1.8): furthest of the {best}: {distance(knn(train, mismatch, best)[1][-1]['z'], mismatch):.2f} sd")
# 4. one real match: its gaps, both forecasts, and its five nearest past matches
r = next(r for r in test if r["match"].startswith("Hearts") and "Hibernian" in r["match"] and r["date"] == "04/10/2025")
p, near = knn(train, r["z"], best)
q = lr(r["z"])
print(f"{r['match']} ({r['date']}), gaps " + ", ".join(f"{v:+.2f}" for v in r["x"]))
print(" KNN:", ", ".join(f"{c} {p[c]:.1%}" for c in RESULTS), "| logistic:", ", ".join(f"{c} {q[c]:.1%}" for c in RESULTS))
for m in near[:5]:
print(f" {m['season'][:2]}/{m['season'][2:]} {m['match']}, gaps " + ", ".join(f"{v:+.2f}" for v in m["x"]))
Further reading
- Nearest Neighbors, scikit-learn user guide: the standard Python library's documentation, covering nearest-neighbour classification and regression, distance weighting and the data structures that make the search fast.
- An Introduction to Statistical Learning, by James, Witten, Hastie and Tibshirani: a free textbook, with KNN and the bias-variance trade-off among its first examples.
- Thomas Cover and Peter Hart, Nearest neighbor pattern classification (IEEE Transactions on Information Theory, 1967): the classic paper on the method and how good it can be.