# Striker or centre-back? Let the simulation decide

Source: https://www.footballdatascience.co.uk/learn/monte-carlo-striker-or-centre-back
Published: 2026-10-02

> When the answer is a chance, you estimate it by playing the season thousands of times. On 2025/26 at halfway, a striker beats a centre-back for the bottom sides and loses for the top ones, and two simple habits stop the simulation's noise from fooling you.

**On the terraces:** Bottom of the league and leaking goals? Everyone says sign a centre-back. This piece plays the rest of the season thousands of times on a computer and finds a striker can be the better buy.

## The football question

Halfway through the 2025/26 Scottish Premiership, **Kilmarnock were 11th, four points above Livingston at the bottom.** The bottom side goes down. Say the board has money for one signing and the manager has two names on the list (made up, like every signing in this piece):

- a **striker** worth 0.15 extra goals a game, about six more over a 38-game season, or
- a **centre-back** worth 0.15 fewer goals conceded a game, about six fewer over a season.

On goal difference they're identical. So does it matter which one? And would the answer be different for Hearts chasing the title, or Hibernian chasing third?

No formula answers that. What the club cares about, the chance of staying off the bottom, depends on the 18 matches still to play, on how good every other side really is, and on luck. So you **simulate it**.

## The concept

[Top at halfway](/learn/simulating-the-league) built a simulation of the rest of a season: a belief about every team's scoring and conceding, from [last season updated with this one](/learn/updating-scoring-rates), and goals drawn from a [Poisson distribution](/learn/poisson-distribution), played out 10,000 times. **Monte Carlo optimisation** puts a decision inside it:

1. **List the options:** the striker, the centre-back, or nobody.
2. **Play the rest of the season thousands of times under each option.**
3. **Count:** the share of runs in which the club reaches its goal is that option's chance.
4. **Pick the option with the highest chance.**

[Shoot, cross or keep the ball?](/learn/expected-utility-shoot-or-cross) scored each choice by the points it was worth on average. Here the score is a chance, and a simulated chance is never exact. Every run is a different season, so the share wobbles, and how much it wobbles is its standard error:

$$\text{standard error} = \sqrt{\frac{p\,(1-p)}{n}}$$

<div class="plain" markdown="1">
In plain football

- **p** is the chance the simulation found, such as 0.75 for staying up in three seasons out of four.
- **n** is the number of simulated seasons.
- **The standard error** is how far the chance could be out through luck alone. With 1,000 runs and p = 0.75 it's 1.4 percentage points; 10,000 runs cut it to 0.43. Ten times the work buys about three times the precision.
</div>

A gap between two options only counts if it's bigger than the [luck margin](/glossary#luck-margin), about two standard errors. When the options are close, that's the whole difficulty.

## A football example

Kilmarnock, the second half of 2025/26 played 10,000 times under each option, the bottom spot to be avoided:

| Kilmarnock | Avoid bottom |
|---|---|
| As they are | 66.4% |
| Sign the striker | 75.7% |
| Sign the centre-back | 73.4% |

Either signing helps a lot. But **the striker is better by 2.2 percentage points** (from the unrounded shares), with a standard error of 0.27, so the gap is about eight times its standard error. A side conceding 1.7 a game should sign a striker?

Here are six clubs, each with its own goal:

<figure class="rank-chart">
<div role="img" aria-label="Striker minus centre-back, in percentage points of each club's chance of its goal. Livingston, avoid bottom: plus 3.3. Kilmarnock, avoid bottom: plus 2.2. Falkirk, top six: plus 0.8. Motherwell, top three: minus 0.3, within the luck margin. Hibernian, top three: minus 1.1. Hearts, the title: minus 1.9.">

</div>
<figcaption>The bottom sides gain more from the striker, the top sides from the centre-back. Only Motherwell's gap is inside its luck margin. 10,000 simulated second halves of 2025/26 for each club, the same runs for both signings.</figcaption>
</figure>

I expected the club's goal to decide it: chasing sides wanting goals, sides protecting a position wanting defenders. It doesn't. What decides it is the club's own numbers: the model's goals scored and conceded a game against an average side (Kilmarnock's 1.04 is about 40 goals over a 38-game season). In bold, the smaller of the two, the end the better signing strengthens (Motherwell's are too close to call):

| Club | For | Against |
|---|---|---|
| Livingston | **1.16** | 1.87 |
| Kilmarnock | **1.04** | 1.72 |
| Falkirk | **1.16** | 1.57 |
| Motherwell | 1.29 | 1.26 |
| Hibernian | 1.64 | **1.21** |
| Hearts | 1.68 | **1.07** |

**Strengthen whichever number is smaller.** For Kilmarnock, 0.15 goals is a 14% lift to their scoring but only a 9% cut in what they let in. For Hearts it's the other way round. In a Poisson match a side's results depend largely on its scoring rate compared with its conceding rate, as a ratio rather than a difference, so the same 0.15 goals goes further at the smaller end.

One match against an average side shows it without any simulation. Kilmarnock's expected points go from 0.927 to **1.034 with the striker** and 1.009 with the centre-back. Hearts' go from 1.788 to 1.881 with the striker and **1.898 with the centre-back**. Over Kilmarnock's 18 remaining games the striker's edge is about half a point, and at the bottom of the table half a point is worth 2.2 percentage points of survival.

### Change the question, change the answer

The 0.15 goals is made up, and so is the idea that a signing adds a fixed number of goals. Suppose instead each signing were worth 15%: the striker lifts scoring by 15%, the centre-back cuts conceding by 15%. Now 15% of Kilmarnock's 1.72 conceded is more goals than 15% of their 1.04 scored, and the answer flips: **centre-back 77.8%, striker 75.7%**, 2.1 points the other way. Hearts flip too, with the striker now 1.4 points better.

So the simulation hasn't found a law of football. It has answered exactly the question it was given. **What the money buys, in goals, is the assumption that decides the answer**, just as the made-up curves decided the best setting in [attacking without losing the defence](/learn/attack-without-losing-the-defence). The simulation's job is to show what that assumption means for the thing the club cares about. Pinning down the assumption is a job for scouting.

## Taming the noise

### The same runs for every option

The gaps above are trustworthy because of one trick. Each run of the season draws every team's strength and a random number for every score. If the striker and the centre-back get separate runs, the gap between them carries two lots of luck. Give them **the same runs**, the same strengths and the same lucky deflections, and the only thing that differs is the signing. This is called **common random numbers**.

For Kilmarnock, the standard error of the gap is 0.27 with the same runs and 0.62 with separate ones. Matching that with separate runs would take about five times as many: (0.62 ÷ 0.27)² is 5.3. And it shows in the decisions. Make Kilmarnock's call 100 times, each time on 1,000 runs: with the same runs for both signings, the striker was chosen **100 times**; with separate runs, **91**. Nine times in a hundred, the club signs the wrong player through simulation luck alone.

One detail makes it work. Each simulated score uses exactly one random number (the `poisson` function in the code below), so the two versions of a run stay in step match by match. The usual way of drawing Poisson goals uses a varying number of random numbers, and the two versions drift apart.

### The optimiser's curse

Now give the manager more choice: split the 0.15 goals between attack and defence in eleven steps, from all defence to all attack. Choosing the best of eleven noisy estimates has a catch. The winner is likely to be the one whose luck ran highest, so its estimate is too good. This is the **optimiser's curse**, a cousin of the winner's curse at auctions.

To measure it, estimate all eleven splits with 500 runs each, pick the best and check it against a fresh 10,000 runs. Ten times over:

| Runs | Too high (points) | Times |
|---|---|---|
| Same | 0.9 | 7&nbsp;of&nbsp;10 |
| Separate | 3.5 | 10&nbsp;of&nbsp;10 |

The fresh runs put every split between 72.5% (all defence) and 74.6% (all attack). With separate runs, the winner's estimate was too high by more than that whole range. The cures are the same runs for every option, and **re-running the winner on fresh runs** before quoting its number.

Even fresh runs wobble. All attack is the same option as the striker above, which came out at 75.7% in its own 10,000 runs; here it's 74.6%. Two separate sets of runs, a gap of 1.1 points, inside the luck margin for comparing them. Noise never goes away; you just measure it.

## Why it matters

- **Most football decisions are scored by a chance:** the title, survival, Europe. Simulation puts a number on each option, even when no formula can.
- **The same runs for every option** cost nothing and here were worth five times the runs.
- **Distrust the winner's number.** Re-run the best option before you believe how good it is.
- **The assumptions decide the answer.** A goal or a percentage? The simulation can't tell you which; it tells you what each would mean.

## Limitations

- **The signings are made up.** A real signing's worth is uncertain too; a fuller version would draw it afresh in every run.
- **The model is part 5's:** goals only, no injuries, form or managers, and no points deductions.
- **A signing counts from the first game after halfway**, with no time to settle in.
- **The fixtures after the split are the real ones**, as in part 5; had results gone differently, some of them would have changed.
- **Avoiding the bottom isn't quite staying up.** In the Scottish Premiership 11th place goes into a play-off; here it counts as safe.
- **One season.** The smaller-number pattern comes from the Poisson model, so expect it elsewhere, but the gaps are for these clubs in 2025/26.

## Try it yourself

Take the dice-and-spreadsheet season from [part 5's try-it-yourself](/learn/simulating-the-league). Play it five times with your side's scoring raised a little, and five times with its conceding lowered by the same amount. Then do it again with the **same** random numbers for both versions (in a spreadsheet, paste them as values so they stop changing). Which comparison is steadier?

## Reproduce the analysis

The results files are published by [football-data.co.uk](https://www.football-data.co.uk/scotlandm.php). Download the Premiership file (SC0) for 2024/25 and 2025/26 and save them as `SC0_2425.csv` and `SC0_2526.csv`; they aren't rehosted on this site. Every run has a fixed seed, so you'll get the same numbers; it takes about four minutes. Then:

```python
import csv
import random
from collections import Counter
from datetime import datetime
from math import exp, factorial, sqrt

def season(s):  # (home, away, home goals, away goals) in date order
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in ("H", "D", "A")]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    return [(r["HomeTeam"], r["AwayTeam"], int(r["FTHG"]), int(r["FTAG"])) for r in games]

def totals(games):  # goals for, goals against and matches, per team
    gf, ga, n = Counter(), Counter(), Counter()
    for h, a, x, y in games:
        gf[h] += x; ga[h] += y; n[h] += 1
        gf[a] += y; ga[a] += x; n[a] += 1
    return gf, ga, n

def table(games, teams):  # points, then goal difference, then goals scored
    pts, gd, gs = Counter({t: 0 for t in teams}), Counter(), Counter()
    for h, a, x, y in games:
        pts[h] += 3 * (x > y) + (x == y); pts[a] += 3 * (y > x) + (x == y)
        gd[h] += x - y; gd[a] += y - x; gs[h] += x; gs[a] += y
    return sorted(teams, key=lambda t: (pts[t], gd[t], gs[t]), reverse=True), pts

def poisson(lam, u):  # one random number per score, so every option gets the same luck
    k, p = 0, exp(-lam)
    total = p
    while u > total:
        k += 1
        p *= lam / k
        total += p
    return k

# Bayesian part 5's model: 2025/26 stopped at halfway, beliefs from last season worth k = 20 matches
games, last = season("2526"), season("2425")
teams = sorted({t for g in games for t in g[:2]})
lf, la, ln = totals(last)
mu = sum(x + y for *_, x, y in last) / (2 * len(last))
home = sum(x for *_, x, _ in last) / len(last) / mu
away = sum(y for *_, y in last) / len(last) / mu
played, rest = games[:len(games) // 2], games[len(games) // 2:]
gf, ga, n = totals(played)
k = 20
pf = {t: lf[t] / ln[t] if ln[t] else 0.85 * mu for t in teams}
pa = {t: la[t] / ln[t] if ln[t] else 1.15 * mu for t in teams}

def chances(club, options, goal, runs, seed):
    """Play the rest of the season runs times; for each option, a list of 1 (club reached its goal) or 0, run by run."""
    hits = {name: [] for name in options}
    for r in range(runs):
        rng = random.Random(seed + r)  # run r is the same season for every option: common random numbers
        att = {t: rng.gammavariate(k * pf[t] + gf[t], 1 / (k + n[t])) for t in teams}
        dfn = {t: rng.gammavariate(k * pa[t] + ga[t], 1 / (k + n[t])) for t in teams}
        luck = [(rng.random(), rng.random()) for _ in rest]
        for name, change in options.items():
            a, d = dict(att), dict(dfn)
            a[club], d[club] = change(a[club], d[club])
            sim = [(h, w, poisson(a[h] * d[w] / mu * home, u), poisson(a[w] * d[h] / mu * away, v))
                   for (h, w, *_), (u, v) in zip(rest, luck)]
            hits[name].append(int(goal(table(played + sim, teams)[0].index(club) + 1)))
    return hits

def share(x):
    return sum(x) / len(x)

def se(x):  # standard error of a share, or of an average difference
    m = share(x)
    return sqrt(sum((v - m) ** 2 for v in x) / (len(x) - 1) / len(x))

# a made-up signing worth 0.15 goals a game against an average side, at one end or the other
SIGNINGS = {"as is": lambda a, d: (a, d), "striker": lambda a, d: (a + 0.15, d), "centre-back": lambda a, d: (a, d - 0.15)}
CLUBS = [("Livingston", "avoid bottom", lambda p: p < 12), ("Kilmarnock", "avoid bottom", lambda p: p < 12),
         ("Falkirk", "top six", lambda p: p <= 6), ("Motherwell", "top three", lambda p: p <= 3),
         ("Hibernian", "top three", lambda p: p <= 3), ("Hearts", "the title", lambda p: p == 1)]
order, pts = table(played, teams)
print("At halfway: position, points, played; the model's goals scored and conceded a game against an average side")
for club, aim, goal in CLUBS:
    print(f"  {order.index(club) + 1:2}. {club}: {pts[club]} pts, {n[club]} played; "
          f"scores {(k * pf[club] + gf[club]) / (k + n[club]):.2f}, concedes {(k * pa[club] + ga[club]) / (k + n[club]):.2f}")

def points(scores, concedes):  # expected points from one match with these goal rates
    p = lambda lam, g: exp(-lam) * lam ** g / factorial(g)
    return sum(p(scores, i) * p(concedes, j) * (3 * (i > j) + (i == j)) for i in range(15) for j in range(15))

print("\nExpected points from one match against an average side on neutral ground")
for club in ("Kilmarnock", "Hearts"):
    a, d = ((k * pf[club] + gf[club]) / (k + n[club]), (k * pa[club] + ga[club]) / (k + n[club]))
    print(f"  {club}: as is {points(a, d):.3f}, " + ", ".join(f"{name} {points(*change(a, d)):.3f}" for name, change in SIGNINGS.items() if name != "as is"))

print("\nChance of reaching the goal, 10,000 runs each, the same runs for every option")
for club, aim, goal in CLUBS:
    hits = chances(club, SIGNINGS, goal, 10000, seed=1)
    gap = [s - c for s, c in zip(hits["striker"], hits["centre-back"])]
    apart = sqrt(se(hits["striker"]) ** 2 + se(hits["centre-back"]) ** 2)
    print(f"  {club} ({aim}): as is {share(hits['as is']):.1%}, striker {share(hits['striker']):.1%}, "
          f"centre-back {share(hits['centre-back']):.1%}; striker minus centre-back {100 * share(gap):+.1f} points, "
          f"standard error {100 * se(gap):.2f} (separate runs: {100 * apart:.2f})")

print("\nIf a signing were worth 15% instead of 0.15 goals")
PERCENT = {"striker": lambda a, d: (a * 1.15, d), "centre-back": lambda a, d: (a, d * 0.85)}
for club, aim, goal in (CLUBS[1], CLUBS[5]):
    hits = chances(club, PERCENT, goal, 10000, seed=1)
    gap = [s - c for s, c in zip(hits["striker"], hits["centre-back"])]
    print(f"  {club} ({aim}): striker {share(hits['striker']):.1%}, centre-back {share(hits['centre-back']):.1%}; "
          f"striker minus centre-back {100 * share(gap):+.1f} points, standard error {100 * se(gap):.2f}")

print("\nKilmarnock, striker or centre-back, decided 100 times on 1,000 runs")
club, aim, goal = CLUBS[1]
same = separate = 0
for i in range(100):
    hits = chances(club, SIGNINGS, goal, 1000, seed=10**6 + 1000 * i)
    same += share(hits["striker"]) > share(hits["centre-back"])
    s = chances(club, {"striker": SIGNINGS["striker"]}, goal, 1000, seed=10**7 + 1000 * i)["striker"]
    c = chances(club, {"centre-back": SIGNINGS["centre-back"]}, goal, 1000, seed=2 * 10**7 + 1000 * i)["centre-back"]
    separate += share(s) > share(c)
print(f"  striker chosen {same} times with the same runs for both, {separate} times with separate runs")

print("\nKilmarnock, the 0.15 goals split eleven ways between attack and defence")
SPLITS = {s / 10: (lambda a, d, s=s / 10: (a + 0.15 * s, d - 0.15 * (1 - s))) for s in range(11)}
fresh = {s: share(x) for s, x in chances(club, SPLITS, goal, 10000, seed=5 * 10**7).items()}
print("  fresh 10,000 runs, share on attack: " + ", ".join(f"{s:.1f} {v:.1%}" for s, v in fresh.items()))
for label, same_runs in (("the same 500 runs for every split", True), ("separate 500 runs for each split", False)):
    over = []
    for i in range(10):
        if same_runs:
            est = {s: share(x) for s, x in chances(club, SPLITS, goal, 500, seed=6 * 10**7 + 500 * i).items()}
        else:
            est = {s: share(chances(club, {s: change}, goal, 500, seed=7 * 10**7 + 500 * (11 * i + j))[s])
                   for j, (s, change) in enumerate(SPLITS.items())}
        pick = max(est, key=est.get)
        over.append(est[pick] - fresh[pick])
    print(f"  {label}: the best-looking split overstated its chance by {100 * sum(over) / 10:+.1f} points on average, "
          f"{sum(o > 0 for o in over)} times of 10")
```

## Further reading

- [Variance reduction](https://en.wikipedia.org/wiki/Variance_reduction), Wikipedia. Ways of getting more precision from fewer simulation runs, common random numbers among them.
- [Stochastic optimization](https://en.wikipedia.org/wiki/Stochastic_optimization), Wikipedia. Optimising when the thing you measure is noisy, including by simulation.
- [Winner's curse](https://en.wikipedia.org/wiki/Winner%27s_curse), Wikipedia. Why the highest bid at an auction tends to be too high: the same effect as the optimiser's curse.
- J. E. Smith and R. L. Winkler (2006), "The Optimizer's Curse: Skepticism and Postdecision Surprise in Decision Analysis", *Management Science*. The paper that named the curse and showed how to correct for it.
