MFL hands every manager an efficiency number for free — the share of his perfect-hindsight ceiling he actually scored. It has never printed the same number for its own projections, and that number is the only honest answer to “should I just have played the chalk.” Same rosters, same weeks, same ceiling: the owners realised 82.00% of it and the projections would have realised 81.86%. Six seasons, and the two of them cannot be told apart — a finding that survives a correction to how the weeks were selected, and gets stronger for it.
Correction — 10 August 2026
This page’s conclusion stands. Its evidence was weaker than stated, and its central p-value was measured on a sample that had been pruned in a way that flattered it.
Every number below was computed on the 1,149 franchise-weeks that pass a coverage gate: keep a week only if at least 90% of its started non-coach slots carry an MFL projection. That gate is defensible on its own terms — the projection cannot start a player it never scored — but it is not neutral. It drops 179 of 1,328 weeks, and it drops them seven times more often from some franchises than others: 21 weeks each off Washington and Las Vegas, three each off West Waco and Willamette. It also drops bad weeks. The removed weeks average 131.60 points for against the survivors’ 161.26. A week fails the gate because the manager started somebody the feed had no forecast for — a deep IDP, a waiver body, an injury fill-in — and those are his worst weeks by construction. So p = 0.409 compares a manager who kept 96% of his record against one who kept 75%, with the missing quarter being the bad quarter.
Re-run without the gate, on the 806 weeks where every started slot carried a projection and the manager and the chalk therefore chose from identical information, the null is stronger, not weaker: permutation p = 0.9093, and the smallest spread of owner career means of any treatment tried. On the players both sides could see, there is no lineup judgment in this league at all. That is now the number to quote, and it says what this page said — on better evidence.
What the gate was deleting turned out to be a real finding, and it is not about judgment. How often a manager starts a player the projection source has never heard of separates franchises cleanly — permutation p = 0.0002, a near four-fold spread from Kawasaki’s 0.976 a week to Willamette’s 0.253 — and it correlates −0.506 with head-to-head wins. It is a roster and engagement trait, not a lineup-judgment one, and it carries its own caveats: a third of those starts were forced, in the sense that no legal lineup of projected players existed on that roster. Full treatment in docs/stranglers/data/owner-dossiers.js.
MFL’s own efficiency is points scored ÷ opt_pts, the lineup a manager would have set with the results in front of him. Nobody can reach it. That does not matter here, because the same impossible ceiling is used as the denominator for both sides — and a shared denominator is exactly what makes a comparison legal even when the level of neither number means anything.
The projection-optimal lineup is built from the same roster, in the same week, with only what was published before kickoff. Take every player the franchise held, rank them by MFL’s own pre-week projection, fill the legal slots, and then score that lineup with what those players actually went on to do. Divide by the same ceiling the manager was divided by. That ratio is the projections’ efficiency score, and it has never been computed.
share of the perfect-hindsight ceiling realised — 1,149 franchise-weeks
The two bars are not a rendering mistake. The gap is +0.14 points of efficiency — fourteen hundredths of one percent of a ceiling neither of them was ever going to reach. In points it is +0.276 a franchise-week to the manager, on a per-week standard deviation of 15.79, which is t = 0.59 and a 95% interval of −0.57 to +0.93. That is not “we could not tell.” It is a precisely bounded nothing: over six seasons, the true edge either way is smaller than one point a week.
So the answer to the question that prompted this is: it would not have mattered. Hand the keys to the projector for six seasons and you lose nothing measurable. Keep them and you gain nothing measurable. Every deviation these sixteen managers made, over 1,149 weeks, netted out to a quarter of a point.
and here is the number that should stop you over-reading the one above
The +0.14 gap is smaller than the arbitrariness in the counterfactual itself. On 195 of the 1,149 weeks — 17.0% — there is more than one lineup tied for projection-optimal, because two players or two whole legal shapes carry the same projected total. Which of them gets called “the projection’s lineup” is a convention. Break every one of those ties in the projection’s favour and its efficiency is 82.18%, ahead of the owners; break every one against it and the number is 81.59%. The convention is worth 0.60 points of efficiency and the measured gap is 0.14, so the owners’ 82.00% sits inside the bracket. That is not a flaw in the measurement. It is the measurement making its own case: the two are close enough that the tie-break rule could decide the winner, which is what a tie means.
This is the whole reason the piece exists in this shape, and the framing came from the league, not from the data.
“Framing it around VP only will taint the numbers through circumstance because maybe they went against experts and were correct but happened to be going against a buzzsaw of an opponent that there was nothing they could do about.”
That is correct and it is not a small point. A Victory Point is a function of the other franchise’s score, which no lineup decision can touch. A manager who benches the projection’s pick, is right by twelve points, and still loses to a 210-point opponent has his correct decision recorded as a defeat. Run enough of those through a VP counterfactual and you are measuring the schedule.
Points For is the clean measure of the choice. It moves if and only if the lineup moved. Everything in this article is PF; the word “win” below never means a head-to-head, only a week in which the manager’s own lineup outscored the projection’s.
The Victory Point counterfactual is a separate measurement and it does not disagree with this one — it finds that the projection reshuffles 111 Victory Points and four of six playoff fields while netting exactly zero, because the scoring rule conserves VP by construction. And the lineup-hour piece reports the identical PF quantity from the other side: −0.28 a week for obeying the projection, over the same 1,149 franchise-weeks. This article is that number with the sign the other way round, +0.276 for the manager, taken apart by franchise, by owner and by week.
Stated before any table, because getting it wrong produces a result just as clean and pointing the other way.
MFL never projects a coach. Not once: 0 of 73,732 projection rows in six seasons carry a coach. But this league starts nineteen players and three of them are coaches, and MFL’s published efficiency — the free one, the one everybody sees — is over all nineteen slots.
So the naive comparison, a 16-slot projection lineup divided by the 19-slot ceiling, hands the manager his three coach slots for free and reports them as judgment. Both sides are therefore re-based on a 16-slot ceiling: the best legal non-coach lineup, picked with hindsight, from the same roster. MFL’s published 19-slot number for these same weeks is 81.96%; the 16-slot version of it is 82.00%, and the projection meets it on identical terms.
| Check | What it proves | Result |
|---|---|---|
| 16-slot ceiling + top 3 coaches vs MFL’s opt_pts | the slot model and the position map are MFL’s | 1 of 1,328 off by more than 0.05 |
| 16 non-coach starters + 3 coach starters vs MFL’s reported score | the starter parse is right | 4 of 1,328 off by more than 0.05 |
Both sets of exceptions are known and named rather than shrugged at. The single ceiling miss is 2022 week 1, Washington Redskin Potatoes, where MFL itself filled only 7 of 8 offensive slots. The four score misses are all the Carbon County Coal Crushers in 2025 weeks 8–11, the four weeks they held id 0836 — a coach MFL lists that season as position “XX” and excludes from its own optimal lineup. Treating him as a coach anyway breaks the ceiling reproduction in exactly those four weeks and nowhere else; the resolver here does what MFL does, which is refuse to classify him.
the gate, and what it costs
A player MFL never projected counts as 0.0 projected, so he is picked last and the projection-optimal lineup silently degrades. The standing guard is projection_coverage_started, and at the same ≥ 0.9 threshold the sibling analyses use it keeps 1,149 of 1,328 franchise-weeks and throws away 179 — 13.5% of the archive, and the price of not measuring absent data. It is not doing the work: ungated, over all 1,328, the owners read 81.79% and the projections 81.76%, a gap of +0.03 instead of +0.14. Same answer, slightly flatter.
The interesting subset was supposed to be the weeks a manager did not field the projection-optimal lineup. It turns out to be nearly all of them.
1,097 of 1,149
95.5% of franchise-weeks differ from the projection-optimal sixteen. In six seasons and sixteen franchises, the lineup a manager set matched the chalk exactly 52 times. “Deviation weeks only” is not a filter; it is the population.
2.80 of 16 slots
The average deviation swaps 2.80 of the sixteen non-coach slots (median 3) and gives up 6.23 projected points doing it. These are not tie-breaks between two players the projection called even.
| Measure | Value |
|---|---|
| Mean per week | +0.289 |
| Median per week | +1.00 |
| Standard deviation | 16.16 |
| t | 0.59 |
| 95% interval on the mean | −0.80 to +1.10 |
| Weeks the deviation gained | 580 |
| Weeks it cost | 500 |
| Dead heats | 17 |
| Six-season total | +317.1 pts |
580 wins against 500 losses — 53.7% of the 1,080 weeks that were not dead heats — and it buys +317.1 points across 1,097 franchise-weeks. Ninety-six franchise-seasons share that total, so it works out at about three and a third points each: roughly a third of what a single starting slot produces in one week, per season. The median deviation week is +1.00, which is the honest picture — a very slight tilt toward the manager, drowned in a per-week spread of sixteen points.
weekly PF delta, manager minus chalk — 10-point bins, orange is a week the projection would have won
The distribution is close to symmetric and it has real tails on both sides. The best single overrule on record is the Enniscorthy Electric Eels in 2022 week 8, +63.0 points over the chalk; the worst is the South Park Cows in 2021 week 13, −65.0. Fourteen weeks came in below −40 and fifteen above +40. Individual weeks are enormous. The average of them is a quarter of a point.
A deviation is not binary and treating it as one throws away the only variable that could carry a signal: how far the manager went. Measured two ways — how many slots he changed, and how many projected points he gave up to change them.
| Slots swapped | n | Mean | Median | t | Weeks won |
|---|---|---|---|---|---|
| 1 | 176 | −0.22 | 0.00 | −0.33 | 86 of 176 |
| 2 | 319 | −0.12 | +0.50 | −0.16 | 164 of 319 |
| 3 | 309 | +2.12 | +2.00 | +2.28 | 180 of 309 |
| 4 | 180 | −0.96 | −0.70 | −0.63 | 90 of 180 |
| 5+ | 113 | −0.81 | +2.00 | −0.37 | 60 of 113 |
| Projected pts given up | n | Mean | t | Weeks won |
|---|---|---|---|---|
| 0–1 | 130 | +1.30 | +2.36 | 50 of 130 |
| 1–3 | 219 | +0.38 | +0.41 | 118 of 219 |
| 3–6 | 307 | +0.10 | +0.12 | 165 of 307 |
| 6–12 | 353 | +0.86 | +0.92 | 188 of 353 |
| 12+ | 126 | −2.14 | −1.09 | 59 of 126 |
Neither ladder is monotone, and that is the finding. If overruling the projection were skill you would expect bigger overrules to pay more; if it were error you would expect them to cost more. What is here instead is a +2.12 bump at exactly three swapped slots, at t = 2.28, with five bins on the board — which is roughly what one expects to see somewhere when nothing is happening. The one band that behaves is the far end: give up twelve or more projected points and the mean turns to −2.14, won in only 59 of 126. Even that is t = −1.09 and cannot be leaned on.
Read together, the two tables say the degree of the deviation does not predict its result. That is the same statement as the headline, arrived at from a direction that could have contradicted it.
The obvious defence of a deviation is that it is informed. A projection published on Tuesday cannot see Friday’s injury report, and starting a player who was Questionable on that final report is separately known to cost around two points a start. Lumping those moves in with whims understates the manager, so they are split out.
Every deviation week is labelled by whether any projection-optimal player the manager left out was carrying a Questionable, Doubtful or Out designation on that week’s official NFL injury report. The join runs through nflverse’s weekly reports and this repo’s gsis crosswalk: 26,987 of 34,630 report rows resolve to an MFL id (77.9%), of which 12,750 carry one of the three flags. From the other direction the coverage is far better — 99.7% of the 45,625 non-coach roster rows in the archive can be resolved to a gsis id — so a missed flag is a gap in the report join, not in the roster.
| Deviation weeks | n | Mean PF delta | t | Weeks won |
|---|---|---|---|---|
| Moved off an injury-flagged player | 301 | +0.901 | +0.87 | 173 of 301 |
| No flagged player involved | 796 | +0.058 | +0.11 | 407 of 796 |
The informed deviations are worth +0.844 more per week than the rest, and that difference is Welch t = 0.72. The sign is where the mechanism says it should be — a manager reacting to a Friday injury report does better than one who simply had a hunch — and the size is not distinguishable from zero. Reporting it either way would be dishonest; it is here as a direction, not a result.
Two limits on it, both real. The flag says the manager moved off a flagged player, not that the flag is why; he may have been chasing a matchup and the injury is coincidence. And 301 of 1,097 is a small slice — most deviations involve nobody on the report at all, which is itself worth knowing, because it means the injury story cannot explain the bulk of what managers do.
The same two efficiencies per franchise, over the gated weeks. Read the n column first — and then read the section after this one, because the spread in this table is smaller than the error bars on it.
| Franchise | n | Owner eff | Chalk eff | Gap | PF / wk | Six-season PF | t | Weeks won |
|---|---|---|---|---|---|---|---|---|
| West Waco Wildcatters | 80 | 82.1% | 80.4% | +1.64 | +3.77 | +302.0 | 1.85 | 47 of 80 |
| Enniscorthy Electric Eels | 76 | 82.3% | 80.7% | +1.59 | +3.16 | +240.5 | 1.74 | 43 of 76 |
| Washington Redskin Potatoes | 62 | 81.3% | 80.1% | +1.25 | +2.48 | +153.9 | 1.14 | 36 of 62 |
| Willamette Valley Coopers | 80 | 82.6% | 81.4% | +1.11 | +2.33 | +186.8 | 1.58 | 48 of 80 |
| Scranton Stranglers | 79 | 82.4% | 81.9% | +0.58 | +1.30 | +102.7 | 0.67 | 43 of 79 |
| San Antonio Stetsons | 75 | 83.2% | 82.6% | +0.56 | +1.11 | +83.0 | 0.76 | 38 of 75 |
| Carbon County Coal Crushers | 77 | 80.9% | 80.6% | +0.25 | +0.50 | +38.5 | 0.33 | 42 of 77 |
| Houston Longhorns | 68 | 84.3% | 84.2% | +0.11 | +0.22 | +14.8 | 0.12 | 32 of 68 |
| San Diego Hopheads | 72 | 83.8% | 84.0% | −0.26 | −0.46 | −32.8 | −0.27 | 38 of 72 |
| South Philly Pigeon Boys | 79 | 82.6% | 83.1% | −0.52 | −0.86 | −68.0 | −0.59 | 35 of 79 |
| Chicago Maesters of the Midway | 71 | 80.8% | 81.5% | −0.74 | −1.46 | −103.5 | −0.72 | 31 of 71 |
| South Park Cows | 72 | 80.9% | 81.7% | −0.76 | −1.58 | −113.7 | −0.60 | 33 of 72 |
| Oeiras Silver Swans | 67 | 83.9% | 84.6% | −0.77 | −1.48 | −98.9 | −0.83 | 26 of 67 |
| Las Vegas Gamblers | 62 | 78.4% | 79.2% | −0.87 | −1.49 | −92.5 | −0.88 | 28 of 62 |
| Hollywood Wookies | 66 | 82.1% | 83.2% | −1.06 | −2.07 | −136.4 | −1.00 | 33 of 66 |
| Kawasaki Samurai | 63 | 79.5% | 80.9% | −1.44 | −2.53 | −159.3 | −1.36 | 27 of 63 |
Eight franchises above the chalk, eight below, and not one of the sixteen reaches t = 2. The widest gap on the board, West Waco at +1.64 points of efficiency, is 80 weeks at +3.77 — and its own 95% band runs from −0.22 to +7.77, which includes zero. The table is real; the ordering in it is not established.
A franchise id is not a person. Between 2020 and 2025 four of the sixteen changed owner, and merging the two tenures under one row would credit one man with another man’s decisions.
| Franchise | First owner | Second owner | Franchise row | What it hides |
|---|---|---|---|---|
| Oeiras Silver Swans | Jason Moore 2020–2024 | Luis Martins 2025 | −1.48 | −2.53 and +3.89 |
| Hollywood Wookies | Patrick Hart 2020 | Doug Yohn 2021–2025 | −2.07 | −1.99 and −2.08 |
| Carbon County Coal Crushers | Justin Elliot 2020–2022 | Paul E Richie Jr 2023–2025 | +0.50 | −1.92 and +2.99 |
| Kawasaki Samurai | Jeff Hartke 2020 | Steven Chrestensen 2021–2025 | −2.53 | +4.47 and −4.17 |
Three of the four handovers move the number materially and one does not. Carbon County reads as a franchise a whisker above the chalk, +0.50 a week; split correctly it is a manager 1.92 points below it followed by a manager 2.99 above. Kawasaki reads as the worst franchise on the board, and the twelve weeks of it that belong to Jeff Hartke are the best per-week figure in the entire study. Hollywood is the exception that proves the rule was still worth applying — two different men, −1.99 and −2.08, essentially the same record.
Co-ownership is not a handover and is not split: Chicago has been Dan and John Paplaczyk throughout, and Scranton has been Kevin Herndon and Nico Martinez throughout. Both are one continuous stint under two names. That leaves 20 owner careers across sixteen franchise ids.
Points For gained or lost against the chalk, per owner stint, with the number of franchise-weeks it rests on and a 95% band on the per-week mean. Sorted by mean, which is not the same as ranked — see the column of bands.
| Owner | Franchise | Seasons | n | Owner eff | Chalk eff | PF / wk | Career ± | 95% band | Weeks won |
|---|---|---|---|---|---|---|---|---|---|
| Jeff Hartke | Kawasaki Samurai | 1 | 12 | 87.3% | 85.0% | +4.47 | +53.6 | −2.50 to +11.43 | 8 of 12 |
| Luis Martins | Oeiras Silver Swans | 1 | 11 | 87.1% | 84.9% | +3.89 | +42.8 | −4.25 to +12.03 | 6 of 11 |
| Chase | West Waco Wildcatters | 6 | 80 | 82.1% | 80.4% | +3.77 | +302.0 | −0.22 to +7.77 | 47 of 80 |
| Mark Kirwan | Enniscorthy Electric Eels | 6 | 76 | 82.3% | 80.7% | +3.16 | +240.5 | −0.39 to +6.72 | 43 of 76 |
| Paul E Richie Jr | Carbon County Coal Crushers | 3 | 38 | 80.5% | 78.9% | +2.99 | +113.5 | −0.40 to +6.38 | 23 of 38 |
| Cyril Handal | Washington Redskin Potatoes | 6 | 62 | 81.3% | 80.1% | +2.48 | +153.9 | −1.77 to +6.73 | 36 of 62 |
| Phil Orr | Willamette Valley Coopers | 6 | 80 | 82.6% | 81.4% | +2.33 | +186.8 | −0.57 to +5.24 | 48 of 80 |
| Kevin Herndon & Nico Martinez | Scranton Stranglers | 6 | 79 | 82.4% | 81.9% | +1.30 | +102.7 | −2.48 to +5.08 | 43 of 79 |
| Chris Perkins | San Antonio Stetsons | 6 | 75 | 83.2% | 82.6% | +1.11 | +83.0 | −1.76 to +3.97 | 38 of 75 |
| Derrick Williams | Houston Longhorns | 6 | 68 | 84.3% | 84.2% | +0.22 | +14.8 | −3.47 to +3.91 | 32 of 68 |
| David Jayanathan | San Diego Hopheads | 6 | 72 | 83.8% | 84.0% | −0.46 | −32.8 | −3.76 to +2.84 | 38 of 72 |
| Jamshaid Muzaffar | South Philly Pigeon Boys | 6 | 79 | 82.6% | 83.1% | −0.86 | −68.0 | −3.73 to +2.01 | 35 of 79 |
| Dan & John Paplaczyk | Chicago Maesters of the Midway | 6 | 71 | 80.8% | 81.5% | −1.46 | −103.5 | −5.44 to +2.53 | 31 of 71 |
| Denver Milam | Las Vegas Gamblers | 6 | 62 | 78.4% | 79.2% | −1.49 | −92.5 | −4.81 to +1.82 | 28 of 62 |
| Chris B | South Park Cows | 6 | 72 | 80.9% | 81.7% | −1.58 | −113.7 | −6.78 to +3.62 | 33 of 72 |
| Justin Elliot | Carbon County Coal Crushers | 3 | 39 | 81.2% | 82.1% | −1.92 | −75.0 | −6.66 to +2.81 | 19 of 39 |
| Patrick Hart | Hollywood Wookies | 1 | 13 | 82.6% | 83.6% | −1.99 | −25.9 | −7.25 to +3.27 | 5 of 13 |
| Doug Yohn | Hollywood Wookies | 5 | 53 | 82.0% | 83.1% | −2.08 | −110.5 | −6.99 to +2.82 | 28 of 53 |
| Jason Moore | Oeiras Silver Swans | 5 | 56 | 83.3% | 84.6% | −2.53 | −141.7 | −6.37 to +1.31 | 20 of 56 |
| Steven Chrestensen | Kawasaki Samurai | 5 | 51 | 77.3% | 79.8% | −4.17 | −212.9 | −8.27 to −0.08 | 19 of 51 |
One of the twenty careers has a 95% band that excludes zero, and it excludes it by eight hundredths of a point. With twenty stints on the board, one is what testing twenty things produces when nothing is there. Steven Chrestensen’s −4.17 over 51 weeks is the largest negative in the study and it is also exactly the result chance was going to hand somebody.
Three of the twenty are single seasons and are printed but not ranked. Jeff Hartke’s 12 weeks, Luis Martins’s 11 and Patrick Hart’s 13 carry bands roughly nine points wide in each direction. They sit at the top and near the bottom of the table for the same reason a coin lands heads three times: there is no career there to read. The two names at the very top of this table are the two names with the least evidence behind them, which is the most useful thing the table does.
If overruling the projection were a skill some managers have, the spread of career means would be wider than the spread you get by dealing the same 1,149 weeks out at random. It is not.
| Quantity | Value |
|---|---|
| Observed sd of owner means | 2.554 |
| Median sd under reshuffling | 2.436 |
| 95th percentile under reshuffling | 3.477 |
| Permutation p | 0.409 |
p = 0.409. The real league is barely more spread out than a league where every manager’s weeks were dealt to him at random, and four times in ten a random deal spreads out further. Whatever separates Chase from Chrestensen in the table above, this measurement cannot show that it is either of them.
And this is the number the correction at the top replaces. 0.409 is measured on the gated 1,149. On the 806 weeks where both sides saw the same men it is 0.9093, with an observed sd of 2.213 against this table’s 2.554 — a tighter league, further from separable, on a cleaner sample. Two other treatments were run and neither is quotable: leaving the gate off but scoring an unprojected player as zero gives 0.052, and it handicaps the chalk, which can then never start a man it has no number for; imputing the position-week median instead gives 0.016, and it over-credits the chalk, which starts an unknown who then scores like the waiver body he is — visible in that treatment’s absurd league mean of +20 points a week against this page’s +0.28. The four treatments disagree because one choice decides the answer, and only the fourth is a fair fight.
The arithmetic behind that says the same thing more usefully. The per-week standard deviation of the PF delta is 15.79. Over the median owner stint that is a standard error of 1.96, so any single owner’s career figure carries a 95% band of ±3.84 points a week — and the entire observed spread, from +4.47 to −4.17, is 8.6 points wide. The error bar on one owner is nearly half the range of all twenty. Pooled over all 1,149 weeks the band tightens to ±0.91, which is why the league-wide statement can be made at all and the per-owner ones cannot.
The shape over a season, and over six of them. Nothing in the headline rests on a single year — but there is one place in the calendar where the tie might not hold, and it is where the mechanism predicts.
| Week | n | PF / wk | t | Weeks won | Owner eff | Chalk eff |
|---|---|---|---|---|---|---|
| 1 | 91 | +3.29 | +2.24 | 50 of 91 | 80.3% | 78.6% |
| 2 | 85 | +0.51 | +0.31 | 46 of 85 | 80.9% | 80.6% |
| 3 | 89 | +1.47 | +0.86 | 50 of 89 | 80.3% | 79.6% |
| 4 | 87 | −1.44 | −0.79 | 37 of 87 | 80.5% | 81.2% |
| 5 | 82 | −0.09 | −0.05 | 43 of 82 | 83.2% | 83.3% |
| 6 | 84 | −1.20 | −0.74 | 41 of 84 | 82.8% | 83.5% |
| 7 | 80 | +0.61 | +0.39 | 41 of 80 | 81.9% | 81.6% |
| 8 | 89 | −1.71 | −0.96 | 39 of 89 | 81.5% | 82.3% |
| 9 | 79 | +0.21 | +0.13 | 39 of 79 | 82.8% | 82.7% |
| 10 | 81 | +1.89 | +1.05 | 46 of 81 | 82.7% | 81.7% |
| 11 | 81 | +0.19 | +0.11 | 43 of 81 | 83.5% | 83.4% |
| 12 | 76 | −1.08 | −0.61 | 32 of 76 | 82.1% | 82.6% |
| 13 | 81 | +0.87 | +0.42 | 43 of 81 | 82.6% | 82.2% |
| 14 | 64 | +0.13 | +0.07 | 30 of 64 | 84.4% | 84.3% |
Both efficiencies climb through the season together — 80.3% and 78.6% in week 1, 84.4% and 84.3% in week 14. Rosters get sorted, injuries resolve, and everybody realises more of their ceiling. They climb in parallel, which is the tie restated.
The exception is the front of the calendar, and it is where a mechanism actually predicts one. A week-1 projection has no in-season data behind it; every later one does. Splitting on that rather than fishing the week table:
weeks of the season · PF a week, manager minus chalk — the lower bar runs right to left
In the first three weeks the managers are ahead by +1.78 a week over 265 franchise-weeks (t = 1.92, owner 80.5% against 79.6%); from week 4 on they are behind by −0.18 over 884 (t = −0.33, 82.5% against 82.6%). The difference between the two is 1.96 points a week at Welch t = 1.83. That is suggestive and it is not established — week 1 on its own is t = 2.24, which sounds better until you notice fourteen weeks were on the board and one of them was going to look like that. The claim worth making is narrow: if there is any window where a manager’s own read beats the chalk, the evidence points at September, and even there it is under two points a week.
| Season | n | Owner eff | Chalk eff | Gap | PF / wk | t | Weeks won |
|---|---|---|---|---|---|---|---|
| 2020 | 190 | 82.4% | 82.2% | +0.18 | +0.35 | +0.31 | 93 of 190 |
| 2021 | 173 | 81.8% | 81.7% | +0.11 | +0.23 | +0.18 | 87 of 173 |
| 2022 | 190 | 81.0% | 81.1% | −0.15 | −0.29 | −0.26 | 92 of 190 |
| 2023 | 188 | 82.3% | 81.5% | +0.80 | +1.59 | +1.43 | 96 of 188 |
| 2024 | 208 | 82.3% | 82.7% | −0.37 | −0.72 | −0.65 | 103 of 208 |
| 2025 | 200 | 82.2% | 81.9% | +0.31 | +0.58 | +0.51 | 109 of 200 |
The largest season on record is 2023 at +1.59 a week, t = 1.43. The smallest is 2024 at −0.72. Six seasons, four positive and two negative, none of them able to carry a claim on its own — and the pooled six carrying one only because pooling is what buys the precision.
Neither efficiency is a grade, and 82.00% is not a score out of 100. opt_pts is chosen after the games are played. Nobody has ever reached it and nobody ever will. The level is meaningless; only the fact that both sides are divided by the same impossible number makes the comparison legal. Reading 82.00% as “the managers are 82% efficient” is the single most likely way to misuse this page.
A deviation is not always a decision. The projection-optimal lineup can include a player nobody would start — a man on injured reserve, or one MFL projected before news broke. The manager who left him out did not overrule an expert; he obeyed reality. The injury split above catches part of that and is honest that it does not catch all of it, and the ≥ 0.9 coverage gate removes the worst of the rest.
The projection-optimal lineup is genuinely ambiguous on one week in six, and the ambiguity is larger than the finding. 195 of the 1,149 gated weeks have more than one lineup tied for projection-optimal; where that happens the luckiest and unluckiest of them are 6.89 points apart on average and as much as 38.5 apart in one week. Bracketed across the whole study that is 81.59% to 82.18%, against the owners’ 82.00%. Any reading of this page that turns +0.14 into “the managers edged it” is reading a number the tie-break rule could have reversed.
An owner with two seasons does not have a career. Every n in the owner table is printed for this reason. The three single-season stints are not ranked, and the seventeen longer ones are ranked only in the weak sense that the numbers are sorted; the permutation test says the sort order carries no information. If you want one number from that table, take the n, not the mean.
This is Points For and it is deliberately blind to whether it won anything. A manager who beat the chalk by twelve and lost his matchup is a +12 here. That is the point — it is what makes the number a measure of his decision — but it means nothing on this page is a claim about standings, playoff odds or Victory Points. The VP question is answered elsewhere and the answer there is that the rule conserves VP, so a league-wide switch to the chalk nets exactly zero no matter what this article had found.
82.00 ÷ 81.86
It is not smarter than the room and the room is not smarter than it. There is no oracle to defer to and no edge to protect by ignoring one. A start/sit argument that ends “but the projection says” has cited a source with a six-season record of a quarter of a point a week.
+20.08 vs +0.276
The sibling measurement puts the whole weekly pass at +20.08 points a week and shows almost all of it is avoiding dead slots — byes, inactives, empty positions. That is where the lineup hour pays. Choosing between two live starters, measured against the chalk here, is worth +0.276.
There is one narrow exception worth holding lightly: September. In weeks 1–3 the projection is running on last season’s data and the managers are ahead by +1.78 a week. It is not significant, it is the one place a mechanism predicts an edge, and it is the only part of this study that would be worth re-testing when 2026 lands. Everywhere else, the honest instruction is that the deliberation is free in both directions — and free is exactly what makes it safe to keep doing.
Every franchise-week in raw/weekly_results for 2020–2025 is rebuilt from the export: the players held, which of them started, what each scored, and MFL’s own opt_pts. The projection-optimal lineup is the best legal 16 non-coach slots by MFL’s pre-week projection from raw/projections — QB 1, RB 2–3, WR 3–4, TE 1–2 filling eight, and DL 2–3, LB 2–3, CB 2–3, S 1–2 filling eight, with DL capped at 2 before 2023 — scored afterwards with what those players actually did. The ceiling is the same slot model run on actual scores. Ties between equally-projected legal shapes are broken on the sorted player-id list, so a rebuild picks the same lineup every time.
Position comes from MFL’s own statement for that season, with a fallback to another season allowed only for non-coach positions, because a coach id in this system is a recycled slot id that carries a different person in different years. Owner identity comes from raw/identity_by_season.json, which records every owner and the seasons each held the franchise; a stint is the unit of a “career” and two owners listed for the same span are co-owners of one stint, not two. Injury flags are nflverse weekly reports joined through this repo’s gsis crosswalk, regular season only.
Four things it cannot see. Weeks 15 onward, because the archive stops at week 14 and 2020 at 13 — so the playoff weeks, the ones that decide anything, are outside this entirely. The three coach slots, which no projection has ever covered and which are a quarter of the fantasy points in this format. Why a manager deviated, beyond the injury flag; there is no record of intent anywhere in the data. And the 179 gated-out franchise-weeks, which are gated out precisely because the projection did not cover them — if the projection is systematically worse where its coverage is thin, this study has removed the evidence for it.
Confidence ceiling 0.49 — realised, not forecast, per the rubric in data/findings/README.md. This is what happened over 1,149 franchise-weeks under MFL’s projections as published between 2020 and 2025. It does not forecast that a better projection would tie as well, and it is not a claim about any projection but MFL’s own.