Skip to content
The Margin
Methodology

AUDIBLE, NFL position ratings

Fit on 2,111 games, 2017 to 2024.

Predicts
Nothing about future games. It describes how strong an NFL team is, from eight position-group ratings.
Inputs
Eight ratings (starting quarterback, receivers, tight end, running back, offensive line, pass defense, rush defense, special teams) fit by ridge regression on 2,111 games from 2017 to 2024 (lambda 19.3, 10,000 bootstraps).
How it is tested
Three checks set before publishing: margin error with the best season left out (13.10 or lower), against the spread with the best season left out (51.5% or higher), and whether all eight weights hold their sign in held-out seasons.
Headline result
Two of three pass: margin error 12.81 and 55.85% against the spread. The sign check fails, 4 of 8 ratings. The composite explains 15.3% of margin variation. Pass defense and the offensive line carry most of it.

Summary

AUDIBLE is a descriptive team-strength composite built from eight position-group ratings, with weights fit by ridge regression. Two of the three checks we set before publishing pass; the third, whether the offensive skill-position weights hold their sign, fails. We publish it anyway because the composite shows how a team is good, even though it has no edge against the closing line.

On 2,111 NFL games from 2017 to 2024, the fitted composite predicts margins about as well as the betting line, inside the noise. But the four offensive skill-position ratings (quarterback, receivers, tight end, running back) add nothing once the closing spread is known: the market already prices them. So AUDIBLE describes how a team is good; it is not a forecast. It is also not the site's NFL team rating, which is the game model's team Elo plus quarterback rating.

1. The strategy pivot

The first plan was to replace the older game model's team inputs by adding up per-position AUDIBLE ratings for each team. An end-to-end test on May 15, 2026 showed no improvement against the spread over the existing model. Adding up the players is useful for describing a team (we can say how a team is good, position by position) but not for forecasting (it doesn't beat the line by enough to call it a betting signal).

This page is the full write-up of that finding: a ridge regression on game results, with 10,000 bootstrap resamples, passed two of the three checks we set in advance but not the third, so we publish it as a description of teams rather than a forecaster.

2. How we got here

v1 (hand-weighted). The first cut used hand-picked proportions: POCKET 0.40, ROUTE 0.20, TRENCH 0.20, HINGE 0.10, ENGINE 0.10, with the defensive unit subtracted at full weight. Calibrated by face-validity, not by data, and overweighted QB relative to what the data actually says (POCKET's no-spread variance share is 9%, not 40%).

Two outside reviews. A statistical review and a football review both predicted the outcome before the fit ran: the line absorbs skill positions efficiently; the “structural” units (OL, defensive sides, ST) are the rock-stable signals.

The fit. We built a table of 2,111 games with the home-minus-away difference in each of the eight ratings plus the closing spread. Leaving one season out at a time chose the ridge penalty (λ = 19.3), and 10,000 bootstrap resamples gave a 95% interval for each weight. Those fitted weights are what this page shows.

3. The eight ratings

4. The fitted weights

We fit two ridge models on the same eight ratings: one with closing spread as a feature, one without. The page publishes the no-spread β as the headline weight (it captures what each rating explains on its own) and reports the with-spread 95% CI for transparency.

STORM-PASSPass defense+2.53[+1.94, +3.12]100%50.9%
TRENCHOL unit+1.29[+0.72, +1.89]100%24.6%
BOOMSpecial teams+1.11[+0.57, +1.67]100%7.1%
STORM-RUSHRush defense+1.01[+0.44, +1.58]100%5.4%
POCKETStarter QB+0.30[-0.29, +0.93]⚠83.5%9.4%
ENGINELead RB+0.22[-0.34, +0.76]⚠77.9%0.7%
ROUTEWR corps+0.19[-0.39, +0.76]⚠73.7%1.9%
HINGETE × target share-0.01[-0.59, +0.57]⚠51.9%0.0%

⚠ = the 95% interval spans zero. Weights are standardized: one standard deviation of improvement in that rating moves the expected home margin by that many points, after controlling for the closing line where applicable).

5. The checks we set before publishing

1. Margin error, best season left out≤ 13.1012.81PASS ✓
2. Against the spread, best season left out≥ 51.5%55.85%PASS ✓
3. Weights hold their signAll 8 ratings: interval excludes zero and the sign never flips in a held-out season4 / 8 (TRENCH, STORM-PASS, STORM-RUSH, BOOM pass; POCKET, ROUTE, HINGE, ENGINE fail)FAIL ✗

Result: the error and spread checks pass; the sign check fails. We publish as a descriptive composite, not as a forecaster. See §7 below for what the failed gate means in football terms.

6. Which ratings carry the composite

Without the spread, the fit explains 15.3% of the variation in game margins on held-out seasons. Each rating's share of that is the most useful read on which signals carry the composite:

STORM-PASS  50.9%  ████████████████████████████████
TRENCH  24.6%  █████████████████
POCKET   9.4%  ██████
BOOM   7.1%  █████
STORM-RUSH   5.4%  ████
ROUTE   1.9%  █
ENGINE   0.7%  ▌
HINGE   0.0% 

STORM-PASS alone explains half of what the eight ratings can say about game outcomes. TRENCH is #2 at 25%. Together, pass defense + offensive-line is three-quarters of the composite's signal. That is what the football review predicted: “defense is more stable than offense, and the pass-D / OL block is the core.”

HINGE (TE × target share) contributes essentially zero to explained variance. ENGINE contributes 0.7%. The original hand-weights gave ENGINE 10% of the composite, that's ~15× overweight. We surface HINGE and ENGINE on the page for descriptive completeness, but the data is clear: at the season-level team-rating grain, RB workload and TE quality add almost nothing on top of OL + defense + ST.

7. Why we don't claim to predict spreads better than the game model

Two separate tests reach the same conclusion. The first ran the player-based team features through the existing spread model and got no measurable improvement. The second, this fit with the spread included, finds the four skill-position ratings have intervals that span zero: the betting market already prices quarterback, receiver, tight end and running back quality, and the marginal information beyond the closing line is approximately zero on the offensive skill side. The composite is no longer shown on team pages, and it is not a betting signal. The site's one NFL team rating is the game model's (see NFL team ratings).

8. Situational signals, tracked, not yet folded in

Alongside the composite, the team drawer surfaces six situational metrics for the season:

These are tracked separately because the football consults call them out as orthogonal to the EPA aggregates already captured in the eight ratings. We'll watch them for a month and decide whether to fold them in. Until then they are explicitly descriptive: shown beside each team but not in the composite.

9. Before and after legal betting spread (2019)

We refit separately on games before 2019 (799) and from 2019 on (1,312), when legal sports betting spread across the US. The closing-spread weight barely moves (0.02). The weights on the individual ratings drift substantially on the offensive side (POCKET, ROUTE, ENGINE, STORM-RUSH all shift by 0.4-0.9 standardized units). The market has changed how it prices these ratings, but its overall calibration is rock-stable.

Practical implication: any production deployment should refit periodically with a rolling window. The bootstrap CIs reported above slightly under-state real uncertainty because the data is non-stationary. The current cadence target is a yearly refit on a rolling 8-year window.

10. Season simulation backtest (May 2026)

We Monte-Carlo simulate each season Y from the prior season's AUDIBLE ratings (Y−1, walk-forward, never any in-season data leakage) and check three forecasting gates on outcomes 2018-2024. Two bugs were caught and fixed during this test:

  1. Defense missing from the 2026 view: the code that builds the composite read the defensive ratings under their old field names after the 2026 data renamed them, so the two defensive ratings (over half of what the composite explains, see section 6) were silently zero on the live site. It now reads both names.
  2. No regression-to-mean between seasons: a single prior-season elite team (BAL 2024 had def-contrib +11.13) was being projected forward at full strength into Y+1. Added a blend of each rating toward the prior-season league mean with a weight of 0.30, picked from five candidates (0.00, 0.15, 0.20, 0.25, 0.30) by the smallest seven-season win-total error.
Win-total average miss (7 seasons)< 2.502.565 → 2.463PASS ✓
Playoff teams picked correctly> 60%53.6% → 71.6%PASS ✓
Super Bowl winner in the top 8≥ 50%4/7 → 4/7PASS ✓

Result (May 15, 2026): all three checks pass. The change that did it was deciding playoff teams separately from win totals: the ridge fit still drives each game's margin and the win totals, and a separate model picks the top seven teams in each conference from the eight ratings, twelve situational stats (third downs, red zone, turnovers, each as a level and a year-over-year change) and coaching changes. The share of playoff teams picked correctly over seven seasons went from 53.6% to 71.6%.

The five weights tried (10,000 simulated seasons each, the winner re-run at 50,000):

0.00  miss 2.565  playoff 53.6%  SB 4/7  → 1 of 3 checks
0.15  miss 2.518  playoff 53.6%  SB 4/7  → 1 of 3
0.20  miss 2.506  playoff 53.6%  SB 4/7  → 1 of 3
0.25  miss 2.497  playoff 53.6%  SB 4/7  → 2 of 3
0.30  miss 2.486  playoff 52.6%  SB 4/7  → 2 of 3 (chosen)

11. Reproducibility

The full report, the 2,111-game training table and the fitted weights are kept with the model's code. The weights in the table above are the ones the team composite reads.

One-line summary

A descriptive composite. It shows which signals it relies on and which it does not. The pass defense and OL unit carry 75% of the signal; everything else is published with a 95% confidence interval that lets the reader judge whether the ⚠ flag applies.