The team that acquires the most does get worse the following season — the correlation is −0.246 and it is real. It is also entirely the record. Hold this season’s win percentage fixed and the effect of acquiring goes to −0.041, which is another way of writing zero. One measure looked like it survived. Every trade carries a date, and the half of it transacted before a single game had been played comes out at −0.020.
Every claim in this piece is a claim about what trading adds on top of knowing the standings. So the standings go first, and they are the number any candidate has to beat.
Win percentage in this league carries forward at r = +0.509. A quarter of next season — 25.9% of the variance — is already written in this season’s record before anybody trades anything. Resampling the sixteen franchises as blocks rather than the eighty pairs as independent draws, the interval is +0.255 to +0.676, width 0.421; the naive pair bootstrap gives +0.314 to +0.661, width 0.346. On this quantity the honest interval is 22% wider than the careless one, and the gap is the size of the mistake available here. It does not go that way on every quantity, which is a correction printed in full in the Method.
That is a strongly mean-reverting league. Three-quarters of next year is not in this year’s number, so any group of teams selected on having had a good season will look like it declined, and any group selected on having had a bad one will look like it improved. Every raw correlation below is contaminated by that before it is contaminated by anything else.
Correlation between a franchise’s trade activity in season t and its win percentage in season t+1. Nothing held fixed.
| Activity in season t | r with win% in t+1 | Reads as |
|---|---|---|
| Players acquired | −0.394 | curse |
| Assets acquired | −0.247 | curse |
| Trades made | −0.157 | curse |
| Distinct partners | −0.090 | curse |
| Picks acquired | −0.077 | flat |
| Picks given up | +0.203 | sell picks, improve |
Six numbers, one clean story: acquire and decline, spend picks and improve. +0.203 is the figure the backlog entry flagged as the surprise — “sell picks, get better,” the reverse of dynasty orthodoxy — and it is the one number that entry quotes in full. It reproduces to the digit, on the sample and the arithmetic it used. The reproduction is not the problem. The interpretation is.
The obvious attack on the table above is mean reversion: teams that buy hard are teams having a good season, and good seasons regress. That attack lands. It lands for the opposite reason than expected.
In this league the acquisitive franchise is the losing franchise. Assets acquired correlates −0.421 with the same season’s win percentage; net assets — everything received minus everything sent — correlates −0.737 with it. The team taking back three pieces for one is the rebuilder at the deadline, not the contender at the trade window.
| Activity in season t | r with win% in the SAME season |
|---|---|
| Net assets (in minus out) | −0.737 |
| Net picks (in minus out) | −0.639 |
| Net players (in minus out) | −0.508 |
| Assets acquired | −0.421 |
| Picks acquired | −0.387 |
| Players acquired | −0.385 |
| Trades made | −0.239 |
| Picks given up | +0.219 |
So the raw −0.247 is not “winners regress.” It is losers persist. The acquisitive teams were bad already, and r = +0.509 says most of a bad season carries over on its own. Same confound, opposite sign, identical consequence: the standings are doing the work, and the only honest question left is what is left over after they have done it.
The same six measures, twice. Above: the raw correlation with next season. Below: the partial correlation, holding this season’s win percentage fixed. Same scale — both ladders run to 0.40 — same order, same eighty pairs. Negative bars run right to left, positive bars left to right.
Raw — correlation with next season’s record
Partial — what is left once the standings are accounted for
The pitched measure — how much a franchise acquires — goes from −0.246 to −0.041. Its 95% interval, resampling franchises, is −0.342 to +0.227. Shuffling which franchise got which season’s activity produces a partial at least that large 74% of the time. On ranks rather than levels it is +0.052 — the other sign, and equally empty. In wins, the point estimate is −0.13 wins per standard deviation of acquisition — that standard deviation is 13.7 assets — over a fourteen-game season, on an interval running from losing 1.16 wins to gaining 0.72.
And picks flip sign. Picks acquired reads −0.110 raw and +0.109 partialled; picks given up drops from +0.222 to +0.131. Both now point the same way, which is the tell that neither is measuring anything: net picks — the difference between the two, the only clean statement of buying versus selling — comes out at +0.025, interval −0.194 to +0.217. The backlog’s “sell picks, get better” does not survive its own confound.
| Measure | same yr | raw next | partial | rank | 95% CI | within-franchise | perm p | family p |
|---|---|---|---|---|---|---|---|---|
| Trades | −0.239 | −0.141 | −0.022 | −0.018 | −0.315 … +0.228 | −0.099 | 0.848 | 1.000 |
| Partners | −0.215 | −0.065 | +0.053 | +0.067 | −0.216 … +0.282 | −0.027 | 0.650 | 0.994 |
| Assets acquired | −0.421 | −0.246 | −0.041 | +0.052 | −0.342 … +0.227 | −0.093 | 0.738 | 0.999 |
| Assets given up | +0.138 | +0.162 | +0.107 | +0.105 | −0.189 … +0.361 | +0.002 | 0.367 | 0.877 |
| Players acquired | −0.385 | −0.376 | −0.227 | −0.109 | −0.464 … +0.042 | −0.235 | 0.049 | 0.251 |
| Players given up | +0.042 | +0.061 | +0.046 | +0.064 | −0.257 … +0.310 | −0.108 | 0.692 | 0.997 |
| Picks acquired | −0.387 | −0.110 | +0.109 | +0.146 | −0.190 … +0.329 | +0.026 | 0.357 | 0.866 |
| Picks given up | +0.219 | +0.222 | +0.131 | +0.076 | −0.127 … +0.345 | +0.113 | 0.258 | 0.764 |
| Net picks | −0.639 | −0.308 | +0.025 | +0.078 | −0.194 … +0.217 | −0.148 | 0.833 | 1.000 |
| Net players | −0.508 | −0.525 | −0.359 | −0.284 | −0.542 … −0.121 | −0.225 | 0.002 | 0.011 |
| Net assets | −0.737 | −0.527 | −0.262 | −0.153 | −0.396 … −0.132 | −0.289 | 0.025 | 0.135 |
Split the eighty pairs into thirds by how much the franchise acquired. Forecast each one’s next season from its record alone — that is the null model, fitted on these same pairs — and compare the forecast to what happened.
| Band | n | Mean assets in | Win% that season | Record-only forecast | Actual next season | Beat the forecast by |
|---|---|---|---|---|---|---|
| Bottom third | 26 | 5.3 | 56.8% | 53.6% | 50.3% | −3.33 |
| Middle third | 28 | 13.0 | 52.2% | 51.1% | 52.8% | +1.67 |
| Top third | 26 | 33.3 | 40.8% | 45.2% | 46.7% | +1.53 |
Read this one as noise, and read the disagreement as the evidence that it is noise. The bands are not monotone — the middle third (+1.67) beats the top (+1.53) — and the continuous partial they are supposed to illustrate is negative (−0.041) while the band slope is positive. Two estimators of one quantity pointing opposite ways at this sample size is exactly what an absent effect looks like, and an earlier version of this article said “the sign is the point.” It is not the point. The size is: the whole spread between the ends is 4.9 points of win percentage, about two-thirds of a win over fourteen games, on 26 pairs a side.
What the table does establish is where the raw correlation came from, and that is the middle column. The most acquisitive third was a 40.8% team; the quietest was a 56.8% team. Sixteen points of win percentage separated them before the year they were being judged on.
Net flow — everything a franchise received minus everything it sent — is the one measure that does not collapse. It is also the one measure whose story changes completely when you ask when the trades happened, which is a question this article did not ask until a reviewer forced it.
Net assets partials to −0.262, interval −0.396 to −0.132, which excludes zero. Demeaning every franchise against its own six-season average gives −0.289. Dropping any single franchise and refitting moves it only between −0.294 and −0.223. Net players, the same axis narrowed to bodies, is stronger still at −0.359 and is the one measure in the family whose permutation p (0.002) survives being judged against the whole family (0.011). In wins: one standard deviation of net acquisition — 8.9 assets, measured on the eighty rows that enter the regression — is worth −1.15 wins the following season, on an interval of −1.83 to −0.54. (An earlier version labelled that standard deviation 8.7, which is the figure over all 96 franchise-seasons; the coefficient it was multiplying comes from the 80-row fit, where the dispersion is 8.89.) By the arithmetic, that is a finding.
Then split it on the date. Every trade in the file carries one. The boundary is the season’s first regular-season NFL game — 10 September 2020 through 4 September 2025, taken from the nflverse schedule already on disk — so one half of a franchise’s net flow is transacted before a single game the season is scored on, and the other half in response to games already played. Across 2020–2025 that is 235 trades before kickoff and 144 after.
| Half of net flow | same yr | raw next | partial | 95% CI | horse race | perm p |
|---|---|---|---|---|---|---|
| Offseason (before kickoff) | −0.497 | −0.268 | −0.020 | −0.265 … +0.229 | −0.049 | 0.863 |
| In-season (after kickoff) | −0.609 | −0.538 | −0.334 | −0.514 … −0.109 | −0.336 | 0.003 |
| Offseason, gross acquisition | −0.312 | −0.056 | +0.125 | −0.157 … +0.379 | +0.149 | 0.291 |
| In-season, gross acquisition | −0.307 | −0.330 | −0.212 | −0.497 … +0.069 | −0.226 | 0.071 |
The half that strictly precedes the season has nothing in it. −0.020, on an interval running from −0.265 to +0.229, with a permutation p of 0.863. The whole of the surviving −0.262 lives in the trades made after the games started going badly — where one standard deviation of net flow, 5.6 assets, is worth −1.25 wins the next season on its own.
This is not a power artifact, which is the first thing to check. The two halves carry the same weight: 1,528 asset-moves offseason against 1,210 in-season inside the pair window, sums of absolute net flow of 278 and 284, and standard deviations of 5.56 and 5.56 assets. They are close to independent, correlating +0.249 with each other. Two comparably sized, comparably dispersed, nearly uncorrelated variables; one carries the entire effect.
Nor is it the boundary. Cutting on the calendar instead — September through November as “in-season,” which moves five early-September trades across the line and makes the in-season share 47.3% of all asset-moves — gives −0.028 offseason and −0.320 in-season. Same answer, one decimal apart.
| Test | Net assets | Offseason half | In-season half |
|---|---|---|---|
| Partial, holding win% fixed | −0.262 | −0.020 | −0.334 |
| Holding Pythagorean expectation fixed instead | −0.240 | +0.035 | −0.327 |
| Holding both fixed | −0.232 | +0.027 | −0.318 |
| Adding last season’s record (n = 64) | −0.322 | −0.043 | −0.417 |
| On ranks rather than levels | −0.153 | +0.013 | −0.200 |
| Within franchise, own six-season mean removed | −0.289 | −0.075 | −0.276 |
| Placebo, pointed at the PREVIOUS season | −0.147 | −0.017 | −0.187 |
An earlier version of this piece named the wrong mechanism, and the table above is why. It said the likely explanation was that win percentage over fourteen games is a noisy reading of team strength, so partialling on it under-removes the confound. That is testable, and it fails. Pythagorean expectation — built from every point scored and allowed, correlating +0.925 with win percentage and forecasting next season marginally better (+0.515 against +0.509) — moves net assets only from −0.262 to −0.240. On both controls, −0.232. Adding the previous season’s record moves it the wrong way. Roughly nine-tenths of the effect survives two independent sharpenings of the control. Noise in the control is not what is happening.
The date is. A franchise’s deadline shape is a readout of a season that has already gone wrong in ways a fourteen-game win percentage under-measures — the injuries, the collapse, the decision to quit, all visible in October and only partly visible in the final record. Selling in October and being bad next year are both consequences of the same October. The backwards placebo agrees: net flow “predicts” the previous season at −0.147, and the in-season half carries almost all of that too (−0.187 against −0.017), which is only possible if the measure is reading team quality rather than causing it.
Two more things weaken it. On ranks — the right check for a long-tailed count — net assets falls about 40%, from −0.262 to −0.153, and the in-season half from −0.334 to −0.200. And the thirds are not monotone: ranked by in-season net flow, the residual against the record-only forecast runs −1.05, +5.94, −5.34, with the middle band beating its forecast by more than either end. Ranked by offseason net flow, the same three numbers are +0.78, −1.44, +0.77 — a flat line, which is what a null is supposed to look like.
The correct summary is narrow and worth having anyway: how a franchise transacts at the deadline carries real information about the following season that its record does not. How it transacts in July carries none. That is a statement about measurement, not about trading, and it is not a decision rule — by the time it is observable, the season it describes is nearly over.
The strongest argument against the headline is that it is underpowered, and the headline is a null. Eighty pairs, sixteen franchises, five seasons each. The smallest partial correlation this design could call significant is 0.218 even under the optimistic assumption that all eighty are independent draws — and they are not. The assets-acquired interval, −0.342 to +0.227, comfortably contains a real buyer’s curse worth a win a season. This piece cannot rule one out. It says the data does not show one, and that the raw correlation everybody would have quoted is fully explained by something else. Those are different claims and only the second is strong.
The strongest argument against the surviving claim is that its treatment is measured at the same time as its control, and partly caused by the same events. An October trade is not a variable that was set before the season and then observed; it is a response to how the season is going, made while the season is still being scored. Holding win percentage fixed does not fix that, because win percentage is the coarser of the two readings. This piece does not resolve the identification problem — it measures how big it is. The component with no such problem, the offseason half, is −0.020. The component drenched in it is −0.334. That is the whole of the effect, and it cuts in the direction of “this is a symptom,” not “this is a cause.”
The date split was a reviewer’s idea, not a pre-registered one. It was run once, on one measure, after that measure had already been identified as the survivor — one degree of freedom, not a new family, which is why the four split measures are held to the original eleven-measure threshold rather than widening it. On that threshold the in-season half still clears, at 0.024. But a test invented after seeing which measure needed explaining deserves a discount, and the honest version of this section is that the split is a strong hypothesis rather than a confirmed one until 2026 is played.
Eleven measures were examined and one beat the family. Net players at family-wise p = 0.011 would be a clean finding if it had been the only question asked. It was not. The family-wise correction is in the table for exactly this reason, and net assets — the more natural quantity, and the one with the tighter interval — does not pass it at 0.135.
Assets are counted, never priced. Six incoming assets could be six waiver-wire flyers or a franchise quarterback and five throw-ins; here they are both a 6. The repo holds contract values and auction prices that would fix this, and deliberately does not use them, because the pitched claim was about volume. A value-weighted version of this question is a different and probably better article.
trades.md flags Hollywood as “high volume plus below-average judgment” and it is right about the judgment. It is not evidence about next year. Trade count partials to −0.022 against next season, interval −0.315 to +0.228, permutation p 0.85 — the emptiest result in the family. A busy counterparty is a counterparty who answers the phone, nothing more.
The measure that looked like it beat the standings only does so from September onward. Before kickoff it is −0.020. The one number worth having for how good a franchise will be next year remains this year’s record, +0.509, a quarter of the variance. That is a low ceiling. It is also the honest one, and it costs nothing to look up.
The 2026 ledger shows why the distinction matters. Through 2 August, sixteen franchises have made 36 trades moving 149 assets. Hollywood is the busiest at fifteen — and a net shedder of fourteen assets, the largest negative flow in the league. San Antonio, with twelve trades, is the largest accumulator at plus fifteen. Read on volume, those two look like the same kind of team; read on net flow, they are opposites. Read correctly, neither reading forecasts anything: every 2026 trade on record falls between 13 February and 2 August, which puts all of it on the offseason side of the split — the side that measures −0.020.
The panel. Sixteen franchises × six seasons = 96 franchise-seasons, 2020–2025, built from raw/weekly_results (83 regular-season week files, 1,328 franchise-weeks, zero ties; 2020 played thirteen weeks and the rest fourteen, which win percentage absorbs). Every season’s league mean win percentage is asserted to be exactly 0.500 before anything is computed — head-to-head play forces it, so a failure there means the aggregate is wrong. Pairing consecutive seasons leaves 80 pairs.
Three trade counts, because the file is bigger than the analysis. raw/trades_all_years.json holds 415 completed trades moving 1,767 assets, classified across the whole file as 853 picks, 837 players and 77 coaches with nothing unclassified, using the same player-index.json rule the metrics store uses. Of those, 379 trades and 1,618 assets fall in the six seasons that have a record; the 36 trades and 149 assets dated 2026 are excluded because that season has not been played. And the activity that actually enters an eighty-pair regression is season t only, 2020–2024: 321 trades, 1,369 assets. The masthead number is the middle one. An earlier version of this article printed 415 in the masthead and 1,767 in this paragraph, both of which describe the file rather than the panel.
Zero-fill, not inner join. The metrics store has no row for a franchise that made no trade — its minimum is one trade, never zero — so joining on presence silently conditions the sample on having traded and drops three franchise-seasons (2020 Washington, 2023 Enniscorthy, 2025 Carbon County). Those are observations of zero activity, not missing data, and are zero-filled here. The restricted 78-pair version is computed alongside and differs trivially: assets acquired partials to −0.051 there against −0.041 here. A correction is owed on this point. An earlier version of this article accused the backlog entry of miscounting — of calling 93 franchise-seasons “with a record” when 96 have one. It does not say that. It says “93 franchise-seasons with a same-season record, 78 with a following-season record,” against a base it declares two lines earlier as 106 franchise-season trade rows. Both figures are exactly right: 106 = 93 rows in 2020–2025 plus 13 franchises with a 2026 trade, and every one of those 93 does have a record. The entry counted trade rows and said so. The zero-fill point stands on its own; the accusation attached to it was a misquotation and is withdrawn.
The estimator. Activity is z-scored within season, because league-wide volume swings from 382 assets moved in 2021 to 217 in 2023 and a raw count otherwise mixes “this franchise was active” with “this was an active year.” Win percentage needs no such treatment. The headline statistic is the partial correlation of activity with next season’s win percentage holding this season’s fixed, equivalently the third coefficient of win%(t+1) ~ 1 + win%(t) + activity(t). p-values are permutation: activity vectors are shuffled among the franchises within a season, which preserves the record autocorrelation and each season’s total volume exactly, 4,000 draws, and the family-wise version records the largest partial anywhere in the eleven measures on each draw. Everything is read from raw files, except the six season-start dates, which come from the nflverse schedule CSV in external/. --verify-db re-checks the build against the metrics store and the two agree on all 744 comparable franchise-season trade values and all 96 records.
A correction on the intervals. All intervals here are a block bootstrap resampling the sixteen franchises, 8,000 draws, seeded. An earlier version of this article said the naive pair bootstrap — which pretends the eighty pairs are independent — “is visibly narrower.” That is true for nine of the eleven measures and for the win% autocorrelation, and false for the survivor it was quoted against. On net assets the clustered interval is −0.396 to −0.132, width 0.265; the naive one is −0.462 to −0.041, width 0.421 — the careful interval is 37% tighter, and the careless one excludes zero by only 0.041. Net picks is the same reversal, 0.411 against 0.527. Clustering is still the right default: it makes no independence assumption that is false here. It is simply not a free width penalty, and the previous sentence implied it always is. Both intervals are now computed and printed for every measure.
Four things it cannot see. Value — assets are counted, not priced. The other three ways a roster changes: the auction, the waiver wire and the rookie draft are all outside this panel, so “trading did not move the needle” is not “roster moves did not move the needle.” Coaches, counted as assets but never analysed separately, 77 of 1,767. And 2026, which has 36 trades on record and no season played yet. Timing used to be on this list and no longer is — it was the thing the list was hiding.
Confidence 0.6 — a null, measured. Eighty pairs is not a sample of this league, it is all of it; there is no more data to get until 2026 is played, which will add sixteen pairs. The result that will hold is the decomposition — that the raw correlation is the standings — because that follows from r = +0.509 and the sign of the same-season correlation, both of which are large and stable. The result that may not is the in-season net-flow effect, which is post-hoc, halves on ranks, non-monotone across thirds, and half-reproduces pointing backwards in time.