# Go for it or play it steady? Planning the run-in backwards

Source: https://www.footballdatascience.co.uk/learn/dynamic-programming-run-in
Published: 2026-10-02

> Bottom of the league with five games to play. Dynamic programming works out how to play every match by starting at the final whistle and working back, and finds that gambling on wins pays when you're two points or more behind, and never when you're level.

**On the terraces:** Bottom of the league with five to play? Some say go for broke, some say keep calm and pick up points. This piece works backwards from the last game to show when each is right.

## The football question

After 33 rounds the Scottish Premiership splits, and the bottom six play each other five more times. Bottom at the split goes down unless it can climb past the side above. Across the 25 seasons with a split since 2000/01, the gap from bottom to 11th at that point ran from nothing to 16 points; it was **4 or fewer in 11 of them**, and the bottom side escaped in **5 of the 25**.

Say you're the manager, a few points behind with five to play. Before every match you choose how to play it:

- **Steady:** win 36.8%, draw 25.2%, lose 38.0%. That's real: it's how the sides in the bottom two at the split actually did in their last five games, 250 matches in all, 1.356 points a game.
- **Bold:** win 42.8%, draw 5.2%, lose 52.0%. Made up: go for the win every time, so more wins, far fewer draws and many more defeats. It's worth slightly *fewer* points on average, 1.336 a game.

The side above you keeps playing steady. **Which way should you play, and does the answer change as the games run out?**

## The concept

**Dynamic programming** solves a chain of decisions by starting at the end. The last game is easy to think about; every game before it leans on the ones after.

1. **Describe every situation, the state:** games left, and points behind.
2. **At the final whistle the answer is known:** ahead of them, you stay up; behind, you go down; level on points, it's goal difference, which we count as a coin toss.
3. **One game earlier,** for each state and each style, every combination of your result and theirs leads to a new state whose chance you already know. Average them, weighted by how likely each combination is, and keep whichever style gives the bigger chance.
4. **Step back one more game and repeat,** until you reach five to play.

It's the backward induction from [who changes first](/learn/sequential-games-substitutions), with one difference: there the other manager was choosing too; here the other side of the coin is luck. The step in the middle is the **Bellman equation**:

$$\begin{aligned} V_n(g) = \max_{\text{style}} \sum \; & P(\text{ours})\,P(\text{theirs}) \\ & \times V_{n-1}(g') \end{aligned}$$

<div class="plain" markdown="1">
In plain football

- **V<sub>n</sub>(g)** is the chance of finishing above them with **n** games left and **g** points behind, playing every remaining game the best way.
- **max over style** means work out steady and bold, and keep the better.
- **The sum** runs over the nine ways a round can go: your win, draw or defeat, times theirs.
- **P(ours) and P(theirs)** are the chances of those results: yours from the style you pick, theirs from their steady record.
- **g′** is the gap after that round: their points minus yours, added on.
- **V<sub>0</sub>** is the final whistle: 1 if you're ahead, 0 if you're behind, a half if you're level.
</div>

The result is a **policy**: not one plan for the run-in, but the right choice for every situation you could find yourself in, so you can look up the answer again after each result.

## A football example

Here is the policy, every state from three points ahead to eight behind, with five games left down to one:

<figure class="rank-chart">
<div role="img" aria-label="The best style for each number of games left and points behind. With three to five games left: steady from three ahead to one behind, bold from two behind or more. With two left: steady at three and two ahead, bold one ahead, steady level, bold one behind or more. With one left: steady at three and two ahead, bold one ahead, steady level and one behind, bold two and three behind; further behind it's already settled.">

</div>
<figcaption>Gold: gamble on the win. Mint: play steady. Grey: the race is already settled, because even winning every game left can't close the gap. Made-up bold style, real steady record.</figcaption>
</figure>

Reading it:

- **Level or ahead: steady,** with one exception below. You're winning the race, and draws keep you winning it. Throwing points away on a gamble makes no sense.
- **Two or more behind: bold, whatever the games left.** You need wins, and a draw barely moves the gap. It's the same logic as being [one down at half-time](/learn/expected-utility-shoot-or-cross): with little to lose, take the gamble.
- **One behind: it depends on the clock.** With three or more to play, steady: a draw is a point to build on. With two left, bold. With one left, steady again, because if they slip up a draw takes you level or above, and a gamble loses half the time.
- **One ahead with two or one to play: bold.** That looks backwards, but work it through for the last game. If they win theirs, you're two behind and only a win saves you; a draw is worth nothing. Working backwards finds corners like this that a rule of thumb never would.

### How much does planning add?

Five games to play, the chance of finishing above them:

| Five to play | The plan | Best habit |
|---|---|---|
| 3 ahead | 76.3% | 76.0% steady |
| 1 ahead | 60.0% | 59.3% steady |
| Level | 50.7% | 50.0% steady |
| 1 behind | 41.5% | 40.7% steady |
| 2 behind | 33.1% | 32.2% bold |
| 4 behind | 18.4% | 18.1% bold |
| 6 behind | 8.5% | 8.3% bold |

The plan beats the best single habit everywhere, but by less than a percentage point. That's honest: the two styles are close on average points, so the choice matters only at the margins. **The wrong habit costs more.** Always bold when one ahead gives 58.1%, almost two points below the plan; always steady when four behind gives 17.2%.

Unlike [striker or centre-back?](/learn/monte-carlo-striker-or-centre-back), there's no simulation noise here. Dynamic programming adds up every possible run of results exactly, so a gap of 0.3 points is a real 0.3 points.

## Why it matters

- **Start at the end.** The last decision is the easy one, and every earlier decision is made easier by knowing it. That's the whole trick.
- **A policy, not a plan.** The right choice depends on where the results have left you, so you need an answer for every situation, not a script.
- **The same match can need different football.** One point behind, the right style changes with the games left; the opponent and the players don't change at all.
- **Small problems, exact answers.** Here there are five games, two styles and a few dozen states. Real clubs have squads, injuries and other rivals, and the states multiply out of reach; that's when simulation, as in [striker or centre-back?](/learn/monte-carlo-striker-or-centre-back), takes over.

## Limitations

- **Bold is made up.** No manager can dial in exact percentages, and how much a gamble really changes wins, draws and defeats would need data this site doesn't have.
- **One rival, played independently.** In reality two or three sides are in the race, and bottom-six sides play each other after the split, so their results are linked.
- **Every match is the same.** Home or away and the strength of the opponent change the real chances game by game.
- **Level on points is a coin toss.** The real tie-break is goal difference, which a side can work on.
- **The steady record rests on 250 games.** Each of its three rates could be out by about six percentage points either way through luck.
- **Points deductions aren't in the data,** so the count of escapes uses results only.

## Try it yourself

Take one game to play, one point behind, and make the numbers easy: steady wins, draws and loses a third of the time each; bold wins half and loses half. Their side wins, draws or loses a third of the time. List the nine ways the round can go for each style, mark which ones finish with you above them (level counts a half), and add up the chances. Which style wins? Now try two points behind.

## Reproduce the analysis

Download the Scottish Premiership files (SC0) for 2000/01 to 2025/26 from [football-data.co.uk](https://www.football-data.co.uk/scotlandm.php), saved as `SC0_0001.csv` and so on; they aren't rehosted on this site. It runs in under a second. Then:

```python
import csv
from collections import Counter, defaultdict
from datetime import datetime
from functools import lru_cache

names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def season(s):  # (home, away, home goals, away goals) in date order
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in ("H", "D", "A")]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    return [(r["HomeTeam"], r["AwayTeam"], int(r["FTHG"]), int(r["FTAG"])) for r in games]

# Each side's 38 games in order: points, goal difference and goals from each. The split comes after 33.
results, before, escaped = Counter(), [], 0
for s in names:
    games = defaultdict(list)
    for h, a, x, y in season(s):
        games[h].append((3 * (x > y) + (x == y), x - y, x))
        games[a].append((3 * (y > x) + (x == y), y - x, y))
    if any(len(g) != 38 for g in games.values()):
        continue  # 2019/20 was stopped by COVID before the split
    rank = lambda upto: sorted(games, key=lambda t: tuple(sum(v[i] for v in games[t][:upto]) for i in range(3)), reverse=True)
    at_split, final = rank(33), rank(38)
    bottom, above = at_split[-1], at_split[-2]
    for t in (bottom, above):  # the bottom two at the split, in their last five games
        results.update(p for p, *_ in games[t][33:])
    gap = sum(p for p, *_ in games[above][:33]) - sum(p for p, *_ in games[bottom][:33])
    before.append(gap)
    escaped += final[-1] != bottom
n = sum(results.values())
STEADY = (results[3] / n, results[1] / n, results[0] / n)  # win, draw, loss
BOLD = (STEADY[0] + 0.06, STEADY[1] - 0.20, STEADY[2] + 0.14)  # made up: more wins, far fewer draws, more defeats
STYLES = {"steady": STEADY, "bold": BOLD}
print(f"Bottom two at the split, last five games, {len(before)} seasons ({n} games): "
      f"win {STEADY[0]:.1%}, draw {STEADY[1]:.1%}, lose {STEADY[2]:.1%}")
for name, (w, d, l) in STYLES.items():
    print(f"  {name}: win {w:.1%}, draw {d:.1%}, lose {l:.1%}, {3 * w + d:.3f} points a game")
print(f"Points from bottom to 11th at the split: {sorted(before)}")
print(f"  4 or fewer behind in {sum(g <= 4 for g in before)} seasons; the bottom side escaped the bottom in {escaped} of {len(before)}")

def finish(gap):  # gap = their points minus ours at the end; level on points is a coin toss on goal difference
    return 1.0 if gap < 0 else 0.5 if gap == 0 else 0.0

def outcomes(style):  # our result and theirs: chance, then the change in the gap
    return [(p * q, theirs - ours) for p, ours in zip(STYLES[style], (3, 1, 0)) for q, theirs in zip(STEADY, (3, 1, 0))]

@lru_cache(maxsize=None)
def best(games_left, gap):
    """The chance of finishing above them from here, playing each game the best way, and that way."""
    if games_left == 0:
        return finish(gap), None
    chance = {style: sum(p * best(games_left - 1, gap + change)[0] for p, change in outcomes(style)) for style in STYLES}
    style = max(chance, key=chance.get)
    return chance[style], style

def habit(style, games_left, gap):  # the same style every game
    if games_left == 0:
        return finish(gap)
    return sum(p * habit(style, games_left - 1, gap + change) for p, change in outcomes(style))

print("\nThe best style by games left (rows) and points behind (columns; minus = ahead). B bold, s steady, . already settled")
print("      " + " ".join(f"{g:>2}" for g in range(-3, 9)))
for games_left in range(5, 0, -1):
    cells = []
    for gap in range(-3, 9):
        chance, style = best(games_left, gap)
        cells.append(" ." if chance in (0.0, 1.0) else " B" if style == "bold" else " s")
    print(f"  {games_left} left" + " ".join(cells))

print("\nChance of finishing above them with five to play")
for gap in (-3, -1, 0, 1, 2, 4, 6):
    where = f"{-gap} ahead" if gap < 0 else "level" if gap == 0 else f"{gap} behind"
    print(f"  {where}: the plan {best(5, gap)[0]:.1%}, always steady {habit('steady', 5, gap):.1%}, always bold {habit('bold', 5, gap):.1%}")
```

## Further reading

- [Dynamic programming](https://en.wikipedia.org/wiki/Dynamic_programming), Wikipedia. Breaking a problem into smaller ones and solving them from the end, with the history of the name.
- [Bellman equation](https://en.wikipedia.org/wiki/Bellman_equation), Wikipedia. The equation above in general form, and the principle of optimality behind it.
- [Markov decision process](https://en.wikipedia.org/wiki/Markov_decision_process), Wikipedia. The general model of decisions made step by step under chance, of which this run-in is a small example.
- N. Hirotsu and M. Wright (2002), "Using a Markov process model of an association football match to determine the optimal timing of substitution and tactical decisions", *Journal of the Operational Research Society*. Dynamic programming inside a single match: when to change tactics or make a substitution.
