Methodology
The platform's rule: separate questions get separate numbers, and every number carries its unit, status, model version, and evidence profile. The table below separates what this release actually publishes from target contracts that remain concept-only, blocked, or otherwise unearned.
Four families, four questions
| Family | Current release | Target contract / blocked scope | Unit |
|---|---|---|---|
| CORE (rate) | Actual is live at research status. B-CORE and role-separated P-CORE describe retrospective realized run value with uncertainty and opportunity counts. | Expected and Forecast remain concept-only target contracts; they are not published products and never share an unlabeled leaderboard with Actual. | Runs above average per role-standard workload: 600 PA, 750 BF for starters, or 250 BF for relievers |
| CROWN (cumulative) | A provisional partial ledger is live. It publishes the supported batting, pitching, standalone baserunning, and replacement lanes. Fielding and catching remain explicit nulls; position and league adjustments are not applied. Residuals remain named rather than redistributed. | The complete-ledger target would require independently supported fielding, catching, position, and league lanes plus public reconciliation. This release does not claim that contract. | Partial runs, converted to partial wins by the published league-season calibration |
| PULSE (form) | Descriptive Actual windows are live at research status. Hitter windows publish realized runs per PA. For pitchers: Each value first averages offensive run value per PA inside an appearance, then gives every appearance equal weight. Lower is better for the pitcher. Windows include both starts and relief outings and disclose the exact mix. Majority role selects only the window lengths: 5 and 10 trailing appearances for a starter majority, or 10 and 20 for a reliever majority; the trailing stream is never filtered by role. Every window retains its exact date, actual opportunity count (PA for hitters; appearances for pitchers), standard error, low-sample warning, and full-season anchor. | This release does not publish a rate deviation or an opponent, park, or role-change decomposition. Role-mix counts disclose composition; they are not a causal role-effect adjustment. Those remain possible target-contract disclosures, not current evidence. | Hitters: realized runs per PA. Pitchers: equal-appearance-weighted mean of within-appearance offensive RV/PA, lower is better |
| SIGNAL (reliability) | Research evidence profiles are live. Six axes — sample, identifiability, granularity, noise, availability, and lineage — remain separate; the weakest load-bearing axis caps the overall band. | SIGNAL does not turn evidence quality into player ability or silently multiply its axes into an estimate. | Per-axis evidence profile and capped overall band |
Registry status
The table above describes what each family publishes. This one is rendered directly from the metric registry — the declarative contract every model must satisfy — so a family’s status here cannot drift from the status the build enforces.
| Family | Question it answers | Status |
|---|---|---|
| CORE | Context-adjusted rate impact per role-standard workload | concept |
| CROWN | Cumulative value above a role-specific replacement baseline | provisional |
| PULSE | Recent form against a full-season baseline | research |
| SIGNAL | How strongly the evidence supports each estimate | research |
Status ladder: concept → scaffold → research → provisional → validated (or blocked). Nothing ships to a default surface below validated without a conspicuous label. A registry family sits at the status its weakest published claim earns, which is why CORE reads lower here than the fitted CORE Actual artifact the ratings route serves.
Published foundations
Base-out run expectancy by season and era (24 states); event run values
(RV = runs scored + RE(after) − RE(before)); RE24; win probability and leverage as
separate context-added lenses; park factors fit only on seasons before the season being adjusted;
era boundaries as versioned contracts (including the 2026 ABS break). Every model run is
replayable from an immutable source capture and a hashed fitted state.
What we refuse to do
No blending of actual and expected value into one score. No "better than WAR" claim without a task-matched, held-out test against declared comparator vintages. No cross-era framing series spanning the 2025 model update and 2026 ABS change. No missing data silently becoming zero. No proprietary metric as a training label. No catcher game-calling value without an identifiable causal design.
Status ladder
concept → scaffold → research → provisional → validated · (blocked / retired)
Promotion requires predeclared mechanical, construct, predictive, and reconciliation gates. Correlation with an existing metric proves similarity, never superiority.