Week 38 ship log
Generated: 2026-09-15 Cadence: as needed (this is a ship log, not an agent audit run) Run type: hand-written from the git log for 2026-09-08 through 2026-09-14 and the plan documents dated the same week. No claims file: nothing here is a pick.
Summary
NFL week 1 graded (the calls sheet went 2-3, -1.76 units at a flat paper unit per call) and the review that followed rebuilt the NFL sheet from four slots to five, replaced the kill bar with one that a fair coin survives, and moved the record's headline from win-loss to wins above market expectation. College football's public power ratings started moving with the season instead of sitting on the frozen preseason number. The NFL player-rating pipeline gained live per-player EPA and in-season updates for every position group. MLB and college baseball jobs were parked for the rest of 2026.
1. McConnell's Calls, NFL sheet: five slots, graded at the number we published (2026-09-14)
Kind: product change, recorded in docs/DECISIONS.md section 13.
Week 1 shipped and graded 2-3 under the August rule (one lock, one dog, one total, two board sides). Three independent review lanes (a bookmaker's view, an analyst's, a data scientist's) looked at it and converged on the same shape, so the sheet was rebuilt:
- Slots. One value dog moneyline (the underdog in the lean band with the largest de-vigged model-minus-market gap), up to three leans on the spread (games where our number sits 1.5 to 3 points from the frozen line, key-number straddles first), and one key-number call (a game whose model number and frozen line sit either side of 3 or 7; labeled experimental). Short weeks ship short; nothing is backfilled from the agree pool or the outliers.
- Retired. The lock (75.6% hit rate, -9.8% ROI on the 2021-2025 replay: a record that measured the market's favorite, not the model), the total (the NFL game card publishes no total, so the sheet contradicted the card), and the closest-to-market board (laying -110 to bet the market's own number).
- Stake. Flat, paper, one unit per call, end to end. The "0.25u tracking stake" language is gone. Week 1's ledger restates from -0.44 to -1.76 units on the stake alone; the calls themselves are untouched.
- The record's headline. Wins above market expectation: actual wins minus the sum of the de-vigged market probabilities of the sides taken. Zero is what no edge looks like whatever the price mix, which a win-loss line cannot say. Week 1: -0.66 with a standard error of 1.09.
- Kill bar. After 100 graded calls, retire or reframe if wins above expectation is below -1.96 SE, or if closing-line value is negative with a 95% interval excluding zero. The bar it replaces (Wilson lower bound under 40% after 36 side calls) fired on a fair coin four times in five.
- Line of record. Bands, sides and gaps are computed from the frozen line captured in the prediction's own provenance, never a line resolved later. On week 1 the request-time line had moved four of fifteen games into a different band and flipped one side.
- Week 1 is not restated under the new rule. It stays 2-3 as published, labeled.
The key-number slot is a pre-registered forward test (docs/plans/NFL_KEY_NUMBER_LEANS_PREREG_2026_09_14.md): on the 2018-2025 panel, leans through 3 or 7 hit 57.5% [49.8, 64.9] on n=160 against 51.0% on n=390 otherwise. One post-hoc cut, not significant, lower bound under the juice. It runs weeks 2 through 18 and nothing about it may be tuned in the window.
Also on /calls: every graded week now renders as a collapsed call table under the season record.
2. NFL: the market detects the news, the page says when we agree with the line (2026-09-14)
Kind: model plumbing, docs/plans/NFL_MARKET_MOVE_NEWS_2026_09_13.md.
ATL @ PIT in week 1 opened PIT -3 in late August and closed -6.5 after the Falcons' quarterback situation became clear. The public card showed -3.0 all week, because the published line was captured at the freeze nineteen days before kickoff, and every channel that could have carried the news was closed: the news pipeline had 29 events pending review with the last review in June, and the QB-aware model was a shadow arm.
What shipped: market moves are captured and read as the detector; news explains a move rather than triggering an adjustment on its own; a ratings adjustment from the move ships behind a gate that only evaluates once a completed game was played under an adjustment that predated its own kickoff (as of this entry, no such game exists, so the gate reads "not evaluable"). Scores now refresh hourly, QB ratings update from games played, and the calls sheet reads the lean band.
3. College football: power ratings move with the season (2026-09-14)
Kind: model change, docs/plans/CFB_SEASON_SIM_RATING_SOURCE_RESULT_2026_09_09.md sections 9.3 and 9.4.
The outside analyst who grades our CSV spot-checks Ohio State's number and had seen 32.332 on every publish since August, because the published rating was the frozen preseason prior. The backtest (6,887 FBS-vs-FBS games, 2016-2025) found a ridge rating fit on completed games, blended with the prior on a ramp, beats the frozen prior by +0.82 margin MAE (95% CI +0.68 to +0.94); the ridge alone is 1.55 MAE worse in weeks 1-3, which is why it is blended rather than swapped.
The blend now runs inside the ratings refresh and every consumer moves together: the rankings API and CSV, the power-rankings archive, the season sim, the championship odds, and the ratings-model game predictions. First application entering week 3: weight 0.26 on 100 games, all 138 ratings moved (Ohio State 32.3 to 30.4, still first; Oregon 27.1 to 20.4; Georgia 28.5 to 25.5). The week-4 checkpoint stands: re-run the 2026 scoring after the Sep 26 games and report whether the season bends toward the ramp. It no longer gates shipping.
Alongside it: the CFB projections CSV now matches Circa lines on the team pair within a day (VSiN dates in ET, the CSV in UTC), a week publishes when it ends, and the week-2 ATS review was written up.
4. NFL player ratings: live EPA, in-season updates, and two bugs (2026-09-09 to 09-10)
Kind: pipeline.
- Per-player EPA is computed in-process from ESPN play-by-play during games, recorded durably, and promoted into the player game-stats table so ratings can run the night a game goes final. The player-stat chain runs automatically when a game finishes.
- CPOE and air yards come from NFL Next Gen Stats rather than nflverse; NGS receiving and rushing snapshots feed RB/WR/TE charting. Volume comes from the box score, which unlocks in-season updates for RB, WR and TE, not only QB.
- Two bugs found and fixed on the way: ESPN clock-stoppage rows were being scored as snaps (inflating EPA), and a pass-attempt bias in the production player models was hiding a units bug. A separate pass closed 95% of the EPA residual and cleared both named suspects.
- nflverse 2026 player weeks are ingested and supersede our live rows when they land; a daily watch (
com.polyedge.nflverse-watch) records availability. - Public site: season-to-date QB production is published (punters no longer served as QBs), in-season player ratings are labeled by the week they run through, a "what happened" narrative renders under each finished game card, and the two QB pages merged into one board at
/nfl/qb.
5. Housekeeping
- MLB and college baseball scheduled jobs are parked (the odds-feed key behind them is dead); the API's startup scan skips MLB while baseball is parked.
- College football week 3 game predictions come from the preseason ratings model under an explicit week-set policy; the page stopped calling a frozen preseason prior "in-season".
- The Gap column on the game board was empty on every row; the rule was fixed rather than the data.
Sources
git log --since=2026-09-08 --oneline(commits 03b48da6 through cee6ddda)docs/DECISIONS.mdsection 13docs/plans/NFL_KEY_NUMBER_LEANS_PREREG_2026_09_14.mddocs/plans/NFL_MARKET_MOVE_NEWS_2026_09_13.mddocs/plans/CFB_SEASON_SIM_RATING_SOURCE_RESULT_2026_09_09.mdoutputs/nfl_weekly_review_panel_2026_09_14/SYNTHESIS.md