# Matches like this one. K-nearest neighbours

Source: https://www.footballdatascience.co.uk/learn/nearest-neighbours
Published: 2026-10-03

> To forecast a match, find the past matches most like it and see what happened. No training, no weights, just the right number of lookalikes. On five SPFL test seasons, the 300 nearest draw level with logistic regression.

**On the terraces:** To call a match, think back to games where the two sides were just as far apart and see how those went. This piece shows that a handful of lookalikes is mostly luck, but a few hundred forecast as well as the best simple model in the series.

## The football question

Before a big match, someone always says "the last time these two met in this kind of form…" and reaches for a memory. **What if, instead of one memory, you looked up every past match that was set up the same way, and counted what happened?**

That's **k-nearest neighbours**, or **KNN**: describe the match, find the *k* past matches most like it, and let them vote.

## The concept

Every model in this series so far has *learned* something: [logistic regression](/learn/logistic-regression) learned weights, a [decision tree](/learn/decision-trees) learned questions. KNN learns nothing. The past matches *are* the model. All it needs is a way to say how alike two matches are, and that's the distance from [measuring player similarity](/learn/player-similarity-distance), with matches in place of players.

Each match is described by the same three numbers as parts 8 to 11, each the home side's figure minus the away side's, in points a game: **recent form**, **this season so far** and **last season**. Two matches are close if all three gaps are close:

$$d = \sqrt{\Delta_1^2 + \Delta_2^2 + \Delta_3^2}$$

<div class="plain" markdown="1">
In plain football

- **Δ₁** is how far apart the two matches' form gaps are, **Δ₂** their this-season gaps, **Δ₃** their last-season gaps.
- Each gap is first measured in standard deviations, so that one feature with bigger numbers can't drown out the others. That's [the scale trap](/learn/player-similarity-distance) again.
- **d = 0** means two matches with exactly the same gaps. The bigger **d**, the less alike they are.
</div>

To forecast a match, KNN measures its distance to every one of the **3,900 training matches**, 2001/02 to 2020/21, takes the **k** nearest, and counts:

$$P(\text{home}) = \frac{H + h}{k + 1}$$

<div class="plain" markdown="1">
In plain football

- **k** is how many neighbours vote: the one setting KNN has.
- **H** is how many of those k neighbours were home wins.
- **h** is the share of home wins across all training matches, about 0.44. It adds one imaginary match, split between home win, draw and away win at their usual rates, so that no result is ever given a chance of zero. A zero would make the [log loss](/learn/model-evaluation) infinite the first time that result happened.
- Draws and away wins are counted the same way, and the three chances add up to 1.
</div>

## One real match

**Hearts 1–0 Hibernian, 4 October 2025.** Before kick-off, Hearts had taken 1.60 points a game more than Hibs over their last five, and 1.33 a game more over the season so far, though Hibs had been 0.16 a game better the season before. Its five nearest training matches:

| Season | Match | Gaps |
|---|---|---|
| 2007/08 | Motherwell 3–0 Gretna | +1.80, +1.29, −0.17 |
| 2012/13 | Hibernian 2–2 Inverness | +1.60, +1.17, −0.16 |
| 2013/14 | Aberdeen 1–0 Ross County | +1.60, +1.13, −0.13 |
| 2002/03 | Kilmarnock 2–2 Aberdeen | +1.80, +1.17, −0.16 |
| 2014/15 | Partick 1–2 St Mirren | +1.40, +1.40, −0.03 |

None of them is an Edinburgh derby, and that's the point: KNN never sees team names, only how far apart the two sides have been. With the 300 nearest voting, it gave Hearts **57.3%**, a draw 21.3% and Hibs 21.4%. Logistic regression said 62.0%, 20.1% and 17.9%.

## How many neighbours?

The whole method rests on **k**, so it has to be chosen honestly: on the held-back seasons, 2016/17 to 2020/21, with the neighbours drawn from 2001/02 to 2015/16. The test seasons stay untouched.

<figure class="rank-chart">
<div role="img" aria-label="Log loss on the held-back seasons against the number of neighbours, from 5 to 800 on a stretched scale. Five neighbours score 1.158, worse than the base rates at 1.073. The score falls to 1.008 at 20 and 0.976 at 100, is best at 300 with 0.9705, then rises slowly to 0.984 at 800. Logistic regression on the same seasons scored 0.968.">

</div>
<figcaption>Log loss on the held-back seasons, lower is better. Too few neighbours and the forecast is mostly luck; too many and it stops being about this match.</figcaption>
</figure>

| Neighbours | Held-back log loss |
|---|---|
| 5 | 1.158 |
| 20 | 1.008 |
| 100 | 0.976 |
| **300** | **0.9705** |
| 800 | 0.984 |
| Base rates | 1.073 |

**Five neighbours do worse than knowing nothing.** Five matches are mostly luck: four home wins among them makes the forecast (4 + 0.44) ÷ 6, a 74% chance of a home win, on evidence you could fit on a bench. That's [too jumpy](/learn/bias-and-variance), high variance, and the same [overfitting](/learn/overfitting) a deep decision tree showed in [part 9](/learn/decision-trees).

**Too many neighbours is too stubborn.** At 800, the neighbours reach further and further from the match being forecast, so every forecast drifts back towards the league average. That's high bias.

In between, **300 is best**. It sounds like a lot, but it's about a tenth of the matches the neighbours were drawn from, and for an evenly matched game all 300 sit within half a standard deviation of it. The curve is fairly flat from 100 to 400, so the exact number matters less than avoiding the extremes.

## Against the rest

Now the one look at the 990 test matches, 2021/22 to 2025/26, with the neighbours drawn from all 3,900 training matches:

| Model | Test accuracy | Test log loss |
|---|---|---|
| Base rates | 47.2% | 1.056 |
| Tree, two questions | 55.5% | 0.954 |
| Forest, groups of 100+ | 54.3% | 0.953 |
| Gradient boosting | 54.5% | 0.951 |
| Logistic regression | 54.8% | 0.950 |
| **KNN, 300 neighbours** | **55.4%** | **0.949** |
| Dixon-Coles | 55.3% | 0.945 |
| Bookmaker | 56.2% | 0.932 |

**KNN draws level with logistic regression.** It's 0.0013 ahead on the test seasons, but the luck margin on that gap is ±0.0065, five times bigger, and on the held-back seasons logistic regression was ahead instead (0.968 against 0.9705). Call it a draw. Like every model before it, KNN never makes a draw the favourite: 663 home wins and 327 away wins picked, no draws.

So the simplest model in the series, one that learns nothing at all, matches the best of the four that do. It's more evidence for what [gradient boosting](/learn/gradient-boosting) found: on these three numbers, the features are the ceiling. Getting past it took a different kind of model, [Dixon-Coles](/learn/dixon-coles-ratings), which rates every team's attack and defence.

### An evenly matched game

Where KNN does stand out is the evenly matched game, every gap 0:

| Model | Home | Draw | Away |
|---|---|---|---|
| Logistic regression | 43.5% | 25.8% | 30.7% |
| **KNN, 300 neighbours** | **37.7%** | **29.3%** | **33.0%** |
| Close matches, training seasons | 38.5% | 28.5% | 33.0% |

The last row, from [part 10](/learn/random-forests), is what actually happened in the 397 training matches where the sides were within 0.2 points a game of each other on both season gaps. KNN lands almost on it, and that's no coincidence: its forecast *is* what happened in nearby matches. Logistic regression has to fit one straight line through every match, so it pushes the home side too high in the middle. KNN only ever looks locally.

## Why it matters

- **Similarity is a model.** "Find me matches like this one" is how recruitment shortlists replacements, how recommendation engines suggest the next thing, and how analysts pick historical comparisons. KNN is that idea, scored honestly.
- **One setting, chosen honestly, is the whole job.** KNN has no weights to learn, only k. Pick it on the test seasons and you'd be fooling yourself; pick it on held-back seasons and the test score means something.
- **Local beats global where the pattern bends.** A straight-line model gets the middle wrong; a model that only asks the neighbours doesn't have to commit to a line.
- **Simple can match clever.** A method that learns nothing drew level with every learner in the series. Try the simple thing first.

## Limitations

- **All features count the same.** Distance gives recent form as much say as the two season gaps, although [logistic regression](/learn/logistic-regression) found it adds almost nothing. Here, dropping it made almost no difference (0.9715 without it against 0.9705 with it, on the held-back seasons), but with a useless feature among many, KNN can't learn to ignore it.
- **Slow to forecast.** Every forecast measures the distance to all 3,900 past matches. Fine here, slow for millions.
- **Thin at the edges.** For an evenly matched game the 300th neighbour is 0.48 standard deviations away. For a big mismatch, gaps of +2.4, +2.0 and +1.8, it's 2.32, nearly five times as far, so the "nearest" 300 include some quite different games.
- **Three features, one league.** No injuries, team news or managers, and only the Scottish Premiership, where two clubs have usually been far ahead of the rest.
- **One held-back split.** k was chosen on one block of five seasons. [Cross-validation](/learn/cross-validation) over several blocks would be steadier.

## Try it yourself

The [Similar Matches Finder](/models/similar-matches) does this live: set the three gaps and it lists the 20 nearest past matches and what happened in them. Set k to 5 and watch the forecast lurch as you nudge a slider; set it back to 300 and it settles.

Or by hand: pick this weekend's most even-looking fixture, then find five past matches between sides that were level on points at the same stage of the season. How many were home wins? Now find twenty. Did the answer settle down?

## Reproduce the analysis

The results files are published by [football-data.co.uk](https://www.football-data.co.uk/scotlandm.php). Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as `SC0_2425.csv`; they aren't rehosted on this site. It takes about half a minute, most of it measuring distances:

```python
import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import exp, log, sqrt

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
RESULTS = "HDA"
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    return games

def points_per_game(games):
    pts, n = Counter(), Counter()
    for r in games:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            n[team] += 1
    return {t: pts[t] / n[t] for t in n}

# the same three features as part 8, each the home side's figure minus the away side's, all known before kick-off
rows = []
for s_last, s in zip(names, names[1:]):
    last = points_per_game(season(s_last))
    promoted = 0.85 * sum(last.values()) / len(last)
    history = defaultdict(list)
    for r in season(s):
        h, a = r["HomeTeam"], r["AwayTeam"]
        if len(history[h]) >= 5 and len(history[a]) >= 5:
            rows.append({"season": s, "match": f"{h} {r['FTHG']}-{r['FTAG']} {a}", "date": r["Date"], "y": r["FTR"], "x": [
                (sum(history[h][-5:]) - sum(history[a][-5:])) / 5,                     # points a game, last five
                sum(history[h]) / len(history[h]) - sum(history[a]) / len(history[a]),  # points a game this season
                last.get(h, promoted) - last.get(a, promoted)]})                        # points a game last season
        for team, p in zip((h, a), POINTS[r["FTR"]]):
            history[team].append(p)

def standardise(fit, other):  # in standard deviations of the fitting matches, so each feature counts the same
    mean = [sum(r["x"][j] for r in fit) / len(fit) for j in range(3)]
    sd = [sqrt(sum((r["x"][j] - mean[j]) ** 2 for r in fit) / len(fit)) for j in range(3)]
    for r in fit + other:
        r["z"] = [(r["x"][j] - mean[j]) / sd[j] for j in range(3)]
    return mean, sd

def distance(a, b):
    return sqrt(sum((p - q) ** 2 for p, q in zip(a, b)))

def knn(fit, z, k, use=slice(0, 3)):  # the share of each result among the k nearest, plus one imaginary match at the base rates
    base = Counter(r["y"] for r in fit)
    near = sorted(fit, key=lambda r: distance(r["z"][use], z[use]))[:k]
    seen = Counter(r["y"] for r in near)
    return {c: (seen[c] + base[c] / len(fit)) / (k + 1) for c in RESULTS}, near

def logistic(fit):  # part 8's model, refitted here so the two can be compared match by match
    def predict(w, z):
        e = {c: exp(w[c][0] + sum(wj * zj for wj, zj in zip(w[c][1:], z))) for c in RESULTS}
        return {c: e[c] / sum(e.values()) for c in RESULTS}
    w = {c: [0.0] * 4 for c in RESULTS}
    for _ in range(200):
        grad = {c: [0.0] * 4 for c in "HA"}
        for r in fit:
            p = predict(w, r["z"])
            for c in "HA":
                miss = p[c] - (c == r["y"])
                grad[c] = [g + miss * v for g, v in zip(grad[c], [1] + r["z"])]
        for c in "HA":
            w[c] = [wj - gj / len(fit) for wj, gj in zip(w[c], grad[c])]
    return lambda z: predict(w, z)

def log_loss(forecasts, data):
    return sum(-log(p[r["y"]]) for p, r in zip(forecasts, data)) / len(data)

def base_rates(fit, data):
    n = Counter(r["y"] for r in fit)
    return log_loss([{c: n[c] / len(fit) for c in RESULTS}] * len(data), data)

# 1. choose k on the held-back seasons 2016/17-2020/21, fitting on 2001/02-2015/16
fit = [r for r in rows if r["season"] < "1617"]
held = [r for r in rows if "1617" <= r["season"] < "2122"]
standardise(fit, held)
print(f"fitting {len(fit)} matches, held back {len(held)}; base rates {base_rates(fit, held):.4f}; "
      f"home wins {sum(r['y'] == 'H' for r in fit) / len(fit):.3f} of the fitting matches")
scores = {}
for k in (5, 10, 20, 50, 100, 200, 300, 400, 600, 800):
    scores[k] = log_loss([knn(fit, r["z"], k)[0] for r in held], held)
    print(f"k = {k:3}: held-back log loss {scores[k]:.4f}")
best = min(scores, key=scores.get)
no_form = log_loss([knn(fit, r["z"], best, use=slice(1, 3))[0] for r in held], held)  # distance on the two season gaps only
print(f"k = {best} without recent form: held-back log loss {no_form:.4f}")
lr_held = logistic(fit)
print(f"chosen: k = {best}; logistic regression on the same seasons {log_loss([lr_held(r['z']) for r in held], held):.4f}")

# 2. refit on all 3,900 training matches and score the 990 test matches once
train = [r for r in rows if r["season"] < "2122"]
test = [r for r in rows if r["season"] >= "2122"]
mean, sd = standardise(train, test)
lr = logistic(train)
p_knn = [knn(train, r["z"], best)[0] for r in test]
p_lr = [lr(r["z"]) for r in test]
picks = Counter(max(RESULTS, key=p.get) for p in p_knn)
right = sum(max(RESULTS, key=p.get) == r["y"] for p, r in zip(p_knn, test))
print(f"test, {len(test)} matches: base rates {base_rates(train, test):.3f}, logistic regression {log_loss(p_lr, test):.4f}, "
      f"KNN {log_loss(p_knn, test):.4f}, KNN accuracy {right / len(test):.1%}, picks {dict(picks)}")
gap = [log(q[r["y"]]) - log(p[r["y"]]) for p, q, r in zip(p_knn, p_lr, test)]  # KNN's log loss minus LR's, match by match
avg = sum(gap) / len(gap)
se = sqrt(sum((g - avg) ** 2 for g in gap) / (len(gap) - 1) / len(gap))
print(f"KNN minus logistic regression: {avg:+.4f}, luck margin +/- {2 * se:.4f}")

# 3. an evenly matched game: every gap 0
even = [(0 - mean[j]) / sd[j] for j in range(3)]
p, near = knn(train, even, best)
q = lr(even)
print("even game, KNN:", ", ".join(f"{c} {p[c]:.1%}" for c in RESULTS), "| logistic:", ", ".join(f"{c} {q[c]:.1%}" for c in RESULTS),
      f"| furthest of the {best}: {distance(near[-1]['z'], even):.2f} sd")
mismatch = [(g - mean[j]) / sd[j] for j, g in enumerate((2.4, 2.0, 1.8))]  # a big mismatch, home side far better
print(f"big mismatch (+2.4, +2.0, +1.8): furthest of the {best}: {distance(knn(train, mismatch, best)[1][-1]['z'], mismatch):.2f} sd")

# 4. one real match: its gaps, both forecasts, and its five nearest past matches
r = next(r for r in test if r["match"].startswith("Hearts") and "Hibernian" in r["match"] and r["date"] == "04/10/2025")
p, near = knn(train, r["z"], best)
q = lr(r["z"])
print(f"{r['match']} ({r['date']}), gaps " + ", ".join(f"{v:+.2f}" for v in r["x"]))
print("  KNN:", ", ".join(f"{c} {p[c]:.1%}" for c in RESULTS), "| logistic:", ", ".join(f"{c} {q[c]:.1%}" for c in RESULTS))
for m in near[:5]:
    print(f"  {m['season'][:2]}/{m['season'][2:]} {m['match']}, gaps " + ", ".join(f"{v:+.2f}" for v in m["x"]))
```

## Further reading

- [Nearest Neighbors, scikit-learn user guide](https://scikit-learn.org/stable/modules/neighbors.html): the standard Python library's documentation, covering nearest-neighbour classification and regression, distance weighting and the data structures that make the search fast.
- [An Introduction to Statistical Learning](https://www.statlearning.com/), by James, Witten, Hastie and Tibshirani: a free textbook, with KNN and the bias-variance trade-off among its first examples.
- Thomas Cover and Peter Hart, *Nearest neighbor pattern classification* (IEEE Transactions on Information Theory, 1967): the classic paper on the method and how good it can be.
