Skip to content

Clue by clue. Naive Bayes, and the danger of counting twice

Start with how often each result happens, then let every clue shift the odds. It's how spam filters work, and on three overlapping football clues it gets worse the more it's told. Merge the clues into one and it draws level with logistic regression.

Intermediate Part 17 of Machine Learning Through Football

New to the notation? The symbols explained

Contents

The football question

Before kick-off you know three things about a match: which side is in better form, which has done better this season, and which did better last season. Each one is a clue. What if you started with how often home sides win, and let each clue nudge the odds in turn?

That's naive Bayes: Bayes' theorem, applied one clue at a time. It's the method behind the first good spam filters, and it's quick, simple and often surprisingly good. On football it has a weakness worth knowing about.

The concept

Start with the base rates: across the training matches, home wins 43.8%, draws 23.4%, away wins 32.8%. Then, for each result, ask how likely the clues are if that result is coming:

$$\begin{aligned} &P(\text{home} \mid \text{clues}) \\ &\quad \propto P(\text{home}) \times P(c_1 \mid \text{home}) \\ &\qquad \times P(c_2 \mid \text{home}) \times P(c_3 \mid \text{home}) \end{aligned}$$

In plain football

  • P(home) is the base rate, how often home sides win before you know anything else.
  • P(c₁ | home) asks: among matches the home side went on to win, how common was a form gap like this one? The same for this season's gap, c₂, and last season's, c₃.
  • ∝ means "in proportion to". Work out the same product for a draw and an away win, then scale the three so they add up to 100%.

How common is a gap like this one? For each result, naive Bayes fits a bell curve, the Normal distribution, to the gaps in matches with that result. Home wins came with a this-season gap of +0.31 points a game on average, give or take 0.74; away wins with −0.39, give or take 0.74. A gap of +0.8 is far more typical of a home win than an away win, so it pushes the odds towards the home side.

The naive part is multiplying the clues together as if each were separate news. That's only fair if, once you know the result, the clues tell you nothing about each other. Spam filters get away with it: thousands of words, each a weak clue, none repeating another much. Football's three gaps are different.

The same three gaps as logistic regression, the same 3,900 training matches, 2001/02 to 2020/21, and the same 990 test matches, 2021/22 to 2025/26.

Clue by clue

Rangers 2–3 Motherwell, 26 April 2026. Rangers had taken 2.40 points a game more than Motherwell over their last five, 0.45 more over the season and 0.68 more the season before. Here's naive Bayes adding one clue at a time. It's the test match where the two versions in this article disagree most, picked to show the problem clearly:

Each clue in turn pushes the naive forecast up, to 91%. One combined strength clue says 61%.

Each step makes sense on its own. Together they don't: the hot form, the better season and the better last season are largely the same news, that Rangers are the stronger side, heard three times. Naive Bayes believes it three times and ends at 91%. Motherwell won.

More clues, worse forecasts

Choose the clues honestly, on the held-back seasons 2016/17 to 2020/21, with the model fitted on 2001/02 to 2015/16:

Clues Held-back log loss
Base rates, no clues 1.073
Recent form only 1.019
Last season only 0.987
This season only 0.983
All three 1.055

Any one clue beats all three together. All three barely beat knowing nothing. The reason is in how alike the clues are: on the training matches, the form and this-season gaps have a correlation of 0.76, this season and last season 0.73, form and last season 0.54. Three clues that overlap that much aren't three pieces of evidence.

The double-counting shows up as overconfidence. In the 160 test matches where all three clues made a home win 80% likely or more, they averaged 92%. Home wins happened 80% of the time.

The fix: one clue, not three

If the clues overlap, merge them. Make one strength clue, a mix of the this-season and last-season gaps, and let naive Bayes use only that. How much of each? Again on the held-back seasons:

This season / last season Held-back log loss
30% / 70% 0.9741
40% / 60% 0.9724
50% / 50% 0.9721
60% / 40% 0.9729
70% / 30% 0.9746

Half and half is best, and better than any single clue. It's the same idea as PCA, which boils overlapping stats down to the few directions that carry most of the information. Recent form is left out: it was the weakest clue alone, and it mostly repeats this season.

Against the rest

The one look at the 990 test matches:

Model Test accuracy Test log loss
Base rates 47.2% 1.056
Naive Bayes, three clues 52.8% 1.020
Naive Bayes, one strength clue 54.6% 0.951
Logistic regression 54.8% 0.950
KNN, 300 neighbours 55.4% 0.949
Dixon-Coles 55.3% 0.945
Bookmaker 56.2% 0.932

With one clue, naive Bayes draws level with logistic regression: 0.0006 behind, with a luck margin of ±0.0041. With three it falls most of the way back to knowing nothing. In the 160 sure-looking matches, the one-clue version averaged 76% against the 80% that happened, slightly cautious rather than badly overconfident.

The three-clue version even picks a draw as the likeliest result 10 times; logistic regression never does. The draw's bell curves are a little narrower than the others (spreads of 0.92, 0.68 and 0.56 points a game against 0.96, 0.74 and 0.61 for home wins), and multiplying three of them exaggerates that for matches near the middle.

Logistic regression doesn't have the overlap problem, because it fits all three weights together: when two features carry the same news, it splits the weight between them rather than counting it twice. That's why its recent-form weight came out close to zero. K-nearest neighbours has the opposite weakness: it can't down-weight a feature at all.

Why it matters

  • Check whether your features say the same thing. Many models assume, somewhere, that features are separate evidence. Correlations between them are quick to check and often big.
  • Overconfidence is the symptom. A model that's too sure, as in calibration in depth, is often counting something twice.
  • Simple models still earn a place. With one well-built clue, a method from spam filtering matches the best simple model in the series, and you can see exactly how every number moves the odds.
  • It shines where clues are many and weak. Text, with thousands of words each nudging the odds a little, is naive Bayes's home ground.

Limitations

  • Bell curves are a choice. The gaps within each result are roughly bell-shaped, but not exactly; the tails, where the mismatches are, fit least well.
  • One combined clue loses detail. Merging the season gaps fixes the double count, but throws away any difference between this season and last that might matter.
  • The fix was chosen on one held-back block. Cross-validation over several blocks would be steadier.
  • Three features, one league. No injuries, team news or managers, and only the Scottish Premiership.

Try it yourself

The Naive Bayes Match Predictor shows the odds moving clue by clue. Set the three gaps, then switch between three clues and one strength clue, and watch the overconfidence appear and disappear.

Or by hand: write down your home-win chance for this weekend's biggest mismatch. Now list every reason you're confident and ask how many of them are really the same reason. Would you knock a few points off?

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It runs in a few seconds:

Show the Python145 lines, ready to copy and run.
import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import exp, log, pi, sqrt
from statistics import correlation

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
RESULTS = "HDA"
CLUES = ["recent form", "this season", "last season"]
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    return games

def points_per_game(games):
    pts, n = Counter(), Counter()
    for r in games:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            n[team] += 1
    return {t: pts[t] / n[t] for t in n}

# the same three gaps as part 8, home side minus away side, in points a game, all known before kick-off
rows = []
for s_last, s in zip(names, names[1:]):
    last = points_per_game(season(s_last))
    promoted = 0.85 * sum(last.values()) / len(last)
    history = defaultdict(list)
    for r in season(s):
        h, a = r["HomeTeam"], r["AwayTeam"]
        if len(history[h]) >= 5 and len(history[a]) >= 5:
            rows.append({"season": s, "match": f"{h} {r['FTHG']}-{r['FTAG']} {a}", "date": r["Date"], "y": r["FTR"], "x": [
                (sum(history[h][-5:]) - sum(history[a][-5:])) / 5,
                sum(history[h]) / len(history[h]) - sum(history[a]) / len(history[a]),
                last.get(h, promoted) - last.get(a, promoted)]})
        for team, p in zip((h, a), POINTS[r["FTR"]]):
            history[team].append(p)

def with_strength(data, w):  # one combined clue: w of this season's gap plus (1 - w) of last season's
    for r in data:
        r["x"] = r["x"][:3] + [w * r["x"][1] + (1 - w) * r["x"][2]]

def naive_bayes(fit, clues):
    """Base rates, then for each clue a bell curve per result: its average and spread among home wins, draws, away wins."""
    prior = Counter(r["y"] for r in fit)
    curves = {}
    for c in RESULTS:
        xs = [r["x"] for r in fit if r["y"] == c]
        for j in clues:
            m = sum(x[j] for x in xs) / len(xs)
            curves[c, j] = (m, sqrt(sum((x[j] - m) ** 2 for x in xs) / len(xs)))
    def forecast(x, upto=None):  # upto: only the first few clues, to watch the odds move clue by clue
        score = {c: log(prior[c] / len(fit)) for c in RESULTS}
        for j in clues[:upto]:
            for c in RESULTS:
                m, sd = curves[c, j]
                score[c] += -log(sd * sqrt(2 * pi)) - (x[j] - m) ** 2 / (2 * sd ** 2)  # log of the bell curve's height
        top = max(score.values())
        e = {c: exp(v - top) for c, v in score.items()}
        return {c: e[c] / sum(e.values()) for c in RESULTS}
    return forecast, prior, curves

def logistic(fit):  # part 8's model, standardised and refitted here, for a match-by-match comparison
    mean = [sum(r["x"][j] for r in fit) / len(fit) for j in range(3)]
    sd = [sqrt(sum((r["x"][j] - mean[j]) ** 2 for r in fit) / len(fit)) for j in range(3)]
    z = lambda x: [(x[j] - mean[j]) / sd[j] for j in range(3)]
    def predict(w, v):
        e = {c: exp(w[c][0] + sum(a * b for a, b in zip(w[c][1:], v))) for c in RESULTS}
        return {c: e[c] / sum(e.values()) for c in RESULTS}
    w = {c: [0.0] * 4 for c in RESULTS}
    data = [(z(r["x"]), r["y"]) for r in fit]
    for _ in range(200):
        grad = {c: [0.0] * 4 for c in "HA"}
        for v, y in data:
            p = predict(w, v)
            for c in "HA":
                grad[c] = [g + (p[c] - (c == y)) * t for g, t in zip(grad[c], [1] + v)]
        for c in "HA":
            w[c] = [a - g / len(data) for a, g in zip(w[c], grad[c])]
    return lambda x: predict(w, z(x))

def log_loss(model, data):
    return sum(-log(model(r["x"])[r["y"]]) for r in data) / len(data)

def base_rates(fit):
    n = Counter(r["y"] for r in fit)
    return lambda x: {c: n[c] / len(fit) for c in RESULTS}

fit = [r for r in rows if r["season"] < "1617"]
held = [r for r in rows if "1617" <= r["season"] < "2122"]
train = [r for r in rows if r["season"] < "2122"]
test = [r for r in rows if r["season"] >= "2122"]
pairs = ((0, 1), (0, 2), (1, 2))
print("how alike the clues are (training matches): " + ", ".join(
    f"{CLUES[i]} and {CLUES[j]} {correlation([r['x'][i] for r in train], [r['x'][j] for r in train]):.2f}" for i, j in pairs))

# 1. on the held-back seasons: which clues, and how to combine the two season gaps
print(f"held back: base rates {log_loss(base_rates(fit), held):.4f}")
for clues in ([0, 1, 2], [0], [1], [2]):
    print(f"  naive Bayes on {', '.join(CLUES[j] for j in clues)}: {log_loss(naive_bayes(fit, clues)[0], held):.4f}")
scores = {}
for w in (0.3, 0.4, 0.5, 0.6, 0.7):
    with_strength(fit + held, w)
    scores[w] = log_loss(naive_bayes(fit, [3])[0], held)
    print(f"  one strength clue, {w:.0%} this season: {scores[w]:.4f}")
best = min(scores, key=scores.get)
print(f"chosen: {best:.0%} this season, {1 - best:.0%} last season")

# 2. refit on all training matches; score the test seasons once
with_strength(rows, best)
naive, prior, curves = naive_bayes(train, [0, 1, 2])
fixed, _, strength = naive_bayes(train, [3])
lr = logistic(train)
for name, model in (("base rates", base_rates(train)), ("naive Bayes, three clues", naive), ("naive Bayes, one strength clue", fixed), ("logistic regression", lr)):
    right = sum(max(RESULTS, key=model(r["x"]).get) == r["y"] for r in test)
    picks = Counter(max(RESULTS, key=model(r["x"]).get) for r in test)
    print(f"test: {name}: log loss {log_loss(model, test):.4f}, accuracy {right / len(test):.1%}, picks {dict(picks)}")
gap = [log(lr(r["x"])[r["y"]]) - log(fixed(r["x"])[r["y"]]) for r in test]
avg = sum(gap) / len(gap)
print(f"one strength clue minus logistic regression: {avg:+.4f}, luck margin +/- {2 * sqrt(sum((g - avg) ** 2 for g in gap) / (len(gap) - 1) / len(gap)):.4f}")

# 3. counting the same news twice: test matches where three clues make a home win 80% or more
sure = [r for r in test if naive(r["x"])["H"] >= 0.8]
print(f"three clues say home 80%+ in {len(sure)} test matches: they average {sum(naive(r['x'])['H'] for r in sure) / len(sure):.0%}, "
      f"one strength clue {sum(fixed(r['x'])['H'] for r in sure) / len(sure):.0%}, home wins happened {sum(r['y'] == 'H' for r in sure) / len(sure):.0%}")
even = [0, 0, 0, 0]
print("even game: three clues", {c: f"{p:.1%}" for c, p in naive(even).items()}, "| one clue", {c: f"{p:.1%}" for c, p in fixed(even).items()})

# 4. clue by clue, for the test match where the two disagree most about a home win
r = max(test, key=lambda r: abs(naive(r["x"])["H"] - fixed(r["x"])["H"]))
print(f"{r['match']} ({r['date']}), gaps " + ", ".join(f"{v:+.2f}" for v in r["x"][:3]) + f", strength {r['x'][3]:+.2f}")
for k in range(4):
    p = naive(r["x"], upto=k)
    print(f"  naive, after {k} clue{'s' if k != 1 else ''}: home {p['H']:.0%}, draw {p['D']:.0%}, away {p['A']:.0%}")
p = fixed(r["x"])
print(f"  one strength clue: home {p['H']:.0%}, draw {p['D']:.0%}, away {p['A']:.0%}")

# 5. what the model learned, for the Naive Bayes Match Predictor
print("base rates:", {c: round(prior[c] / len(train), 4) for c in RESULTS})
for j, label in enumerate(CLUES + ["strength"]):
    source = strength if j == 3 else curves
    print(f"{label}:", {c: tuple(round(v, 4) for v in source[c, j]) for c in RESULTS})

Further reading

  • A Plan for Spam, Paul Graham, August 2002: the essay that popularised Bayesian spam filtering, with the word-by-word probabilities it combines.
  • Naive Bayes, scikit-learn user guide: the standard Python library's documentation, including the Gaussian version used here.
  • An Introduction to Statistical Learning, by James, Witten, Hastie and Tibshirani: a free textbook whose classification chapter covers naive Bayes alongside logistic regression and KNN.

Get the weekly email

One email a week, on Fridays, with everything new, and the occasional note. Unsubscribe in one click. How your email is used.