Methodology
The platform's rule: separate questions get separate numbers, and every number carries its unit, status, model version, and evidence profile. The table below separates what this release actually publishes from target contracts that remain concept-only, blocked, or otherwise unearned.
CORE — Rate impact
Publication state: Research.
- What does it measure?
- How much better or worse than an average player this player was, per role-standard workload.
- What does it read?
- Every plate appearance in the completed 2025 season, converted to run value against the base-out state it happened in, then shrunk toward the league mean by a tournament-selected estimator.
- How should it be read?
- As a rate, never a total. Hitters, starters and relievers sit on three different denominators — 600 PA, 750 batters faced, 250 batters faced — and are three separate questions. The interval matters: where it covers zero, the estimator cannot separate this player from league average.
- What should it not be taken to mean?
- That it is WAR, or comparable to it. That a hitter’s number can be ranked against a pitcher’s. That park, defence or opponent quality have been adjusted for — none of the three is applied in this edition.
- Where is the exact evidence?
- Open CORE for the 2025 run-value leaders, with the sealed-holdout findings beneath them. Every figure there carries the artifact it came from under “Inspect the machinery”.
CROWN — Cumulative value
Publication state: Provisional.
- What does it measure?
- What a player was worth in total across every part of the game this release can measure, above a replacement-level baseline for their role.
- What does it read?
- The component ledger: batting, pitching, standalone baserunning and replacement lanes, each booked from the same event corpus, converted to wins by a published league-season calibration.
- How should it be read?
- As a partial ledger, because it is one. Each component has its own column, so what is missing for a given player is visible on that player’s own row.
- What should it not be taken to mean?
- That it is complete. Fielding and catching are not measured and appear as missing with a stated reason, never as zero. Pitching still contains team defence. No positional or league adjustment is applied, and no residual is redistributed to make a total look finished.
- Where is the exact evidence?
- Open CROWN for the partial ledger, what is missing from it, and the calibration. Every figure there carries the artifact it came from under “Inspect the machinery”.
PULSE — Recent form
Publication state: Research.
- What does it measure?
- How a player’s last stretch compares with that same player’s own full season.
- What does it read?
- Trailing windows over the same event corpus — 50 and 100 plate appearances for a hitter, 5 and 10 or 10 and 20 appearances for a pitcher — each carrying its exact date, opportunity count and standard error.
- How should it be read?
- As a record of what happened. Every window sits beside that player’s own season anchor, so form is never read against the league by accident. For a pitcher the number runs the other way: lower is better.
- What should it not be taken to mean?
- That it forecasts anything. A chronological forward test found these short windows have higher error than assuming league average, and that test is published on the page itself.
- Where is the exact evidence?
- Open PULSE for the form boards and the forward test that failed. Every figure there carries the artifact it came from under “Inspect the machinery”.
SIGNAL — Evidence quality
Publication state: Research.
- What does it measure?
- How strongly the evidence supports a published construct — never how good a player is.
- What does it read?
- Six axes kept separate: sample, identifiability, granularity, noise, availability and lineage.
- How should it be read?
- As a cap, not a score. The weakest load-bearing axis sets the overall band, and the axis that limits it is always named.
- What should it not be taken to mean?
- That a weak band means a player is weak. It is a statement about what the data can support, and it never multiplies its axes into an estimate.
- Where is the exact evidence?
- Open SIGNAL for each construct’s axes and the axis that limits it. Every figure there carries the artifact it came from under “Inspect the machinery”.
Four families, four questions
The comparison view. Each family’s own section is above: CORE, CROWN, PULSE, SIGNAL.
| Family | Current release | Target contract / blocked scope | Unit |
|---|---|---|---|
| CORE (rate) | Actual is live at research status. B-CORE and role-separated P-CORE describe retrospective realized run value with uncertainty and opportunity counts. | Expected and Forecast remain concept-only target contracts; they are not published products and never share an unlabeled leaderboard with Actual. | Runs above average per role-standard workload: 600 PA, 750 BF for starters, or 250 BF for relievers |
| CROWN (cumulative) | A provisional partial ledger is live. It publishes the supported batting, pitching, standalone baserunning, and replacement lanes. Fielding and catching are not measured in this edition; position and league adjustments are not applied. Residuals remain named rather than redistributed. | The complete-ledger target would require independently supported fielding, catching, position, and league lanes plus public reconciliation. This release does not claim that contract. | Partial runs, converted to partial wins by the published league-season calibration |
| PULSE (form) | Descriptive Actual windows are live at research status. Hitter windows publish realized runs per PA. For pitchers: Each value first averages offensive run value per PA inside an appearance, then gives every appearance equal weight. Lower is better for the pitcher. Windows include both starts and relief outings and disclose the exact mix. Majority role selects only the window lengths: 5 and 10 trailing appearances for a starter majority, or 10 and 20 for a reliever majority; the trailing stream is never filtered by role. Every window retains its exact date, actual opportunity count (PA for hitters; appearances for pitchers), standard error, low-sample warning, and full-season anchor. | This release does not publish a rate deviation or an opponent, park, or role-change decomposition. Role-mix counts disclose composition; they are not a causal role-effect adjustment. Those remain possible target-contract disclosures, not current evidence. | Hitters: realized runs per PA. Pitchers: equal-appearance-weighted mean of within-appearance offensive RV/PA, lower is better |
| SIGNAL (reliability) | Research evidence profiles are live. Six axes — sample, identifiability, granularity, noise, availability, and lineage — remain separate; the weakest load-bearing axis caps the overall band. | SIGNAL does not turn evidence quality into player ability or silently multiply its axes into an estimate. | Per-axis evidence profile and capped overall band |
Registry status
The table above describes what each family publishes. This one is rendered directly from the metric registry — the declarative contract every model must satisfy — so a family’s status here cannot drift from the status the build enforces.
| Family | Question it answers | Status |
|---|---|---|
| CORE | Context-adjusted rate impact per role-standard workload | concept |
| CROWN | Cumulative value above a role-specific replacement baseline | provisional |
| PULSE | Recent form against a full-season baseline | research |
| SIGNAL | How strongly the evidence supports each estimate | research |
Status ladder: concept → scaffold → research → provisional → validated (or blocked). Nothing ships to a default surface below validated without a conspicuous label. A registry family sits at the status its weakest published claim earns, which is why CORE reads lower here than the fitted CORE Actual artifact the ratings route serves.
Published foundations
Base-out run expectancy by season and era (24 states); event run values
(RV = runs scored + RE(after) − RE(before)); RE24; win probability and leverage as
separate context-added lenses; park factors fit only on seasons before the season being adjusted;
era boundaries as versioned contracts (including the 2026 ABS break). Every model run is
replayable from an immutable source capture and a hashed fitted state.
What we refuse to do
No blending of actual and expected value into one score. No "better than WAR" claim without a task-matched, held-out test against declared comparator vintages. No cross-era framing series spanning the 2025 model update and 2026 ABS change. No missing data silently becoming zero. No proprietary metric as a training label. No catcher game-calling value without an identifiable causal design.
Status ladder
concept → scaffold → research → provisional → validated · (blocked / retired)
Promotion requires predeclared mechanical, construct, predictive, and reconciliation gates. Correlation with an existing metric proves similarity, never superiority.