Two measurements off the same six seasons look like they contradict each other. Leaving last week’s lineup alone costs twenty points a week. Obeying the projection instead of your own read is worth nothing at all. Both are true, and once you see where the twenty points come from they stop being in tension — because they were never measuring the same act.
The counterfactual is exact, not modelled. Take a franchise’s week W−1 starters, start those same nineteen players in week W, and score them with week W’s real scores. Restricted to the weeks where every one of last week’s starters is still on the roster, so nothing has to be substituted in and no rule has to be invented.
The managed lineup beats the stale one by +20.08 points a week — 11.5% of a 174.23-point week, more than a full starting slot. The measurement is not fragile: sd 23.15, t = +24.3 across 786 comparable franchise-weeks of the 1,232 that exist. The active lineup wins in 644 of them, loses 116 and ties 26. The median week is +16.75.
Weekly gap, active minus stale — 10-point bins, orange is a week doing nothing would have won
A manager changes 4.38 of his nineteen slots in an average week and carries 14.62 of them over untouched. Those four-odd changes are what the twenty points are paid for.
The same archive says something that sounds like the opposite: that the weekly lineup is not a lever at all.
| Counterfactual | What it replaces | n | Mean / wk | t |
|---|---|---|---|---|
| Start last week’s lineup again | the whole weekly pass | 786 | +20.08 | +24.3 |
| Start the projection’s best lineup | the manager’s judgment only | 1,149 | −0.28 | −0.59 |
Obeying the projection every week, over 1,149 franchise-weeks where the projection covered at least 90% of the started lineup, would have returned −0.28 points a week — a precisely bounded nothing. Managers’ departures from the projection are not errors. They are free.
So one measurement says the weekly pass is worth twenty points and the other says the weekly decision is worth zero. The reconciliation is that the pass and the decision are different acts, and only one of them is where the money is.
A slot is dead if it scored nothing — either the player never took the field, or he played and produced 0.0. Counted per franchise-week, in the same 786 weeks.
| Dead slots per week | Stale lineup | Live lineup | Difference |
|---|---|---|---|
| Never took the field | 2.41 | 0.43 | +1.99 |
| — of which, on a bye | 1.49 | 0.106 | +1.38 |
| Played, scored 0.0 | 1.07 | 1.15 | −0.08 |
| All dead slots | 3.48 | 1.58 | +1.90 |
The stale lineup fields 3.48 dead slots a week; the managed one fields 1.58. Almost the whole difference is players who never suited up — 2.41 against 0.43. The two lineups take the same bad luck: 1.07 slots a week where a starter played and scored nothing, against the manager’s 1.15. He is no better at dodging a bad game than a lineup nobody looked at.
And 1.49 of the stale lineup’s dead slots each week are simply players on a bye. The managed lineup carries 0.106 — roughly one bye starter every ten weeks, against one and a half every single week.
9.94 vs 10.00
Points per slot that scored anything at all, pooled over all 786 weeks: 9.94 for the stale lineup, 10.00 for the managed one. Six hundredths of a point. Whatever the manager is doing, it is not picking better players.
1.90 × 10.00
1.90 extra dead slots a week, valued at what a live slot is actually worth, is 19.0 of the 20.08-point gap. The rest of the weekly pass — every genuine start/sit call — splits the remainder.
Split the 786 weeks by how many more dead slots the stale lineup would have fielded than the live one did. If the twenty points are hygiene, the gap should be a function of that number and nothing else.
extra dead slots the stale lineup would have started — and what the weekly pass was worth
It is close to a straight line, and it starts at nothing. Every additional dead slot the manager cleared is worth about ten points — which is exactly what a live slot scores.
In the 177 weeks where the stale lineup was no deader than the live one, active management was worth +1.52 points a week — t = 1.22, indistinguishable from zero, and won in 82 of 177. That is the same nothing the projection measurement found, arrived at from the other direction. Take the dead slots out and the lineup hour is worth roughly a point and a half.
Bye weeks are published before the season starts. Splitting on how many of last week’s starters are on a bye this week uses nothing the manager did not already know on Sunday morning.
| Byes in the stale lineup | n | Gap / wk | Active better |
|---|---|---|---|
| 0 | 351 | +10.05 | 242 of 351 |
| 1 | 108 | +17.41 | 92 of 108 |
| 2 | 117 | +24.31 | 110 of 117 |
| 3 | 104 | +31.12 | 98 of 104 |
| 4+ | 106 | +40.47 | 102 of 106 |
In 435 of the 786 weeks at least one of last week’s starters was on a bye, and in those weeks doing nothing costs 28.16 points. In the 351 weeks where none were, it still costs 10.05 — that residue is injuries and inactives, which are on the injury report rather than the schedule but are just as knowable before kickoff.
The managers in this league are already good at this. Across all 786 weeks, a bye-week player was left in a starting lineup on 60 occasions — 7.6% of weeks. The twenty points are not a habit anybody here has; they are the size of the hole the habit is plugging.
“Doing nothing” is not a strategy anyone in this league plays. Nobody has ever fielded last week’s lineup twice on purpose. So +20.08 is not a rival’s exploitable habit and not a number to expect to gain — it is the size of the hole that a lineup pass is already filling, which is a different and much less exciting claim.
The split that carries the reconciliation conditions on realised zeros. Whether a slot ended up dead is known Sunday night, not Sunday morning, so “+1.52 when the dead slots match” decomposes what happened; it does not promise that a manager who makes only judgment calls gains nothing. The version built purely on what was knowable — the bye ladder above, and the weeks before the first NFL bye, where the gap is +5.39 (t = 3.68, n = 133) — puts the judgment residue small but not at zero. Somewhere between one and five points a week is the honest range for the deliberating part of the hour.
The 786 weeks are a selected two-thirds. The other 446 are excluded because a starter had been dropped, leaving nothing to re-start. That selection is not neutral — but it runs against the finding, not for it. Score a dropped starter as the zero he would actually be and the gap over all 1,232 franchise-weeks is +25.26 (t = 35.4), larger than the number reported here.
A dead slot is not always a mistake. Some of the 1.58 the manager still fields are bench-worse-than-the-starter weeks with no fix available, which is why the target is never zero.
The gap, the dead-slot counts and the direction, season by season. Nothing here rests on a single year.
| Season | n | Gap / wk | Active better | Dead slots, stale | Dead slots, live |
|---|---|---|---|---|---|
| 2020 | 105 | +20.8 | 86 of 105 | 3.49 | 1.17 |
| 2021 | 136 | +21.1 | 117 of 136 | 3.54 | 1.51 |
| 2022 | 133 | +18.2 | 109 of 133 | 3.29 | 1.51 |
| 2023 | 141 | +23.1 | 117 of 141 | 3.70 | 1.76 |
| 2024 | 133 | +20.6 | 114 of 133 | 3.53 | 1.82 |
| 2025 | 138 | +16.8 | 101 of 138 | 3.34 | 1.59 |
Six seasons, six positive means, a spread of 16.8 to 23.1, and the stale lineup fields between 3.29 and 3.70 dead slots in every one of them. This is a property of a nineteen-slot lineup meeting the NFL bye schedule, not a run of luck.
byes · inactives · empty slots
Nineteen slots, one question each: is this man playing this week? That pass captures roughly nineteen of the twenty points and takes minutes. Do it before anything else, and do it even in a week you have no time to think.
worth ~1.5 pts
Once no dead slot is left, choosing between two live starters is worth about a point and a half a week — and the projection is worth −0.28, so there is no oracle to defer to either. The deliberation is the cheap half of the hour. Waivers, trades and the coach slots are where the rest of it should go.
It also sizes a build. An automatic stale-lineup alert — “you have a starter on a bye” — is worth the difference between catching those slots and not, which the bye ladder puts at about seven points per bye slot. It is worth writing, and it does not need to be clever, because the expensive failure it prevents is not a judgment error.
The counterfactual takes each franchise’s starters list from week W−1 and scores exactly those player ids with their week-W scores from the same archive file. A franchise-week is comparable only if every one of last week’s starters is still on that franchise’s week-W roster — 786 of 1,232 — which removes any need for a substitution rule and any freedom to choose one. Consecutive weeks within a season only; nothing crosses a season boundary. The reconstructed live score matches MFL’s own reported franchise score in all 786 weeks, which is the check that the starter parsing is right.
A dead slot is a scored zero, and that is a slightly blunt instrument. A player with no score row at all (never active) and a player who played and finished on 0.0 are both dead, and the table above splits them because they mean different things. Bye is resolved against the real NFL schedule — the teams absent from that week’s fixture list — joined to players by their team in that season’s player file. That team is end-of-season, so a player traded mid-season can be misattributed; the size of that error is visible and tiny, since 1.450 of the 1.49 bye slots have no score row and only 0.013 a week scored points despite being marked on a bye.
Three things it cannot see. Weeks 15 onward, because the archive stops at week 14 (2020 at week 13), so the playoff weeks — the ones that decide anything — are outside this entirely. Whether the manager’s bench could have done better still, which is a different question with a different answer. And the counterfactual’s own realism: a manager on genuine autopilot would also stop making waiver claims, so his roster would decay in ways this holds fixed.
Confidence ceiling 0.49 — realised, not forecast, per the rubric in data/findings/README.md. This is what happened to sixteen franchises over six seasons under this lineup format. Change the number of starting slots or the bye structure and the size of the prize changes with it; the mechanism would not.