Tactical Reroll
⋯

The baseline battery: list share, points and frozen round 11 on every panel (2026-10-03)

From reports/tiers/2026-10-03-baselines.md , rendered when the site is built.

T-570, step 1 of design/tier-plan.md's queue (section 10). Built: tools/tierbaseline.js, a new script beside tools/tierbench.js (the workbench reads the Tyranid raw table; the battery reads only docs/data/rankings.json and round 11's published JSON, so it is its own file). Command: node tools/tierbaseline.js --json reports/tiers/2026-10-03-baselines.json (under a second; 2,000 bootstrap resamples, 2,000 shuffles, 200 random tie-breaks, fixed seeds 570 / 5700 / 57000). Round 11 is read from reports/tiers/2026-10-02-tyranids-round11.json, not re-run: its ρ (0.66, 0.14, −0.15), its within-one counts on the reviewers' shapes (36, 24, 38) and its per-army median (0.173) reproduce exactly. No build/units.json and no BSData checkout needed. Our numbers only; lists counted from grimstat-corpus (CC BY 4.0). A report: nothing on the site changes.

In short. Round 11 is near zero on the panels it was not tuned on. Pooled over the two held-out panels (Space Marines and Orks, 75 units) its ρ is −0.03 [−0.27, 0.20]; list share alone gets 0.21 [−0.03, 0.44] and points alone about 0. Round 11 beats list share on none of the three panels, and beats points (the better sign) only on the Tyranids, the panel we tuned on. Its per-army median ρ with list share is 0.173, with 10 of 27 armies under 0.10 and 5 negative. By the plan's rule this result says redesign the jobs (the matrix and the allocation rule, T-568 and T-569) before fitting any weights: a fitter would be tuning a composite that does not yet order units the way reviewers do outside one army.

1. The headline: pooled held-out ρ

Held-out: the panels no round was tuned on (Tista Minis' Space Marines, Sprues & Brews' Orks). Pooled ρ: each panel's ρ weighted by its units; the interval resamples units within each panel (2,000 draws). "All panels" adds Auspex's Tyranids, which steered rounds 5 to 11, so it flatters round 11.

PredictorHeld-out pooled ρ [95%] (75 units)All three panels [95%] (119 units)
List share alone0.21 [−0.03, 0.44]0.43 [0.28, 0.57]
Points alone, cheaper is better0.01 [−0.21, 0.24]−0.03 [−0.22, 0.16]
Points alone, dearer is better−0.01 [−0.24, 0.21]0.03 [−0.16, 0.22]
Frozen round 11−0.03 [−0.27, 0.20]0.23 [0.05, 0.39]
Shuffle0.000.00

"Held-out" is generous: the Marine and Ork panels have been reported every round since round 6, and one round-11 fix (the loadout bug) was found by looking at a Marine unit. They were never the target of a fit or a switch choice, which is what the label means here. The two locked panels Jordan supplies will be the first truly unseen ones.

2. Every panel, with lift over its shuffle

ρ: Spearman between the predictor and the panel's letters, with a bootstrap 95% interval over the panel's units. Lift: the predictor's number minus the mean of 2,000 random orders of the same units. p: the share of shuffles that reach the predictor's ρ. Within one and same: our order over the panel's graded units, cut at the panel's own letter counts (the reviewer's shape); ties in a predictor (list share ties many units at 0) are broken at random and averaged over 200 draws. Every panel unit matched a unit in the data and in round 11's pool: nothing unmatched, nothing dropped.

Space Marines (Tista Minis, 30 units, shape 4/11/9/6/0; held out)

Shuffle: ρ 0.00 [−0.37, 0.37]; within one 21.6 (97.5th percentile 26); same 8.4.

Predictorρ [95%]LiftpWithin oneLiftSameLift
List share alone0.07 [−0.30, 0.40]+0.070.3522.0+0.47.5−1.0
Points, cheaper is better0.34 [−0.06, 0.67]+0.340.0327.2+5.612.6+4.1
Points, dearer is better−0.34 [−0.67, 0.06]−0.340.9717.5−4.16.5−1.9
Frozen round 110.14 [−0.22, 0.49]+0.140.2324.0+2.49.0+0.6

Orks (Sprues & Brews, 45 units, shape 3/13/25/3/1; held out)

Shuffle: ρ 0.00 [−0.28, 0.29]; within one 37.7 (97.5th percentile 41); same 18.1.

Predictorρ [95%]LiftpWithin oneLiftSameLift
List share alone0.30 [−0.03, 0.58]+0.300.0239.1+1.422.0+3.9
Points, cheaper is better−0.20 [−0.48, 0.10]−0.200.9136.0−1.717.1−1.0
Points, dearer is better0.20 [−0.10, 0.48]+0.200.0940.0+2.319.3+1.2
Frozen round 11−0.15 [−0.44, 0.17]−0.150.8438.0+0.317.0−1.1

The worked example of tier-engine.md 4.5, measured: on this lumpy shape a random order puts 37.7 of 45 within one, so round 11's 38 is +0.3 over chance.

Tyranids (Auspex Tactics, 44 units, shape 13/7/13/4/7; tuned on)

Shuffle: ρ −0.01 [−0.31, 0.30]; within one 22.1 (97.5th percentile 28); same 10.3.

Predictorρ [95%]LiftpWithin oneLiftSameLift
List share alone0.81 [0.64, 0.91]+0.82<0.00141.4+19.325.8+15.5
Points, cheaper is better−0.11 [−0.45, 0.24]−0.110.7723.0+0.914.5+4.2
Points, dearer is better0.11 [−0.24, 0.45]+0.120.2128.0+5.918.0+7.7
Frozen round 110.66 [0.42, 0.82]+0.67<0.00136.0+13.920.0+9.7

Auspex ranks "what is used to success now", so list share at 0.81 is close to the panel's own basis: a reference point, not a cap (plan section 2). List share alone already clears Jordan's 85% line on this table (41.4 of 44 within one, the line is 38).

3. The per-predictor table

PredictorSpace Marines ρ (lift)Orks ρ (lift)Tyranids ρ (lift)Held-out pooledWithin one, lift: SM / Orks / Tyr
List share alone0.07 (+0.07)0.30 (+0.30)0.81 (+0.82)0.21 [−0.03, 0.44]+0.4 / +1.4 / +19.3
Points, cheaper is better0.34 (+0.34)−0.20 (−0.20)−0.11 (−0.11)0.01 [−0.21, 0.24]+5.6 / −1.7 / +0.9
Points, dearer is better−0.34 (−0.34)0.20 (+0.20)0.11 (+0.12)−0.01 [−0.24, 0.21]−4.1 / +2.3 / +5.9
Frozen round 110.14 (+0.14)−0.15 (−0.15)0.66 (+0.67)−0.03 [−0.27, 0.20]+2.4 / +0.3 / +13.9
Shuffle0.000.00−0.010.000

Points' sign. It flips by panel: cheaper is better for Tista Minis' Marines, dearer is better for the Orks and, weakly, the Tyranids. Pooled with one fixed sign, points are worth nothing; choosing the sign per panel after seeing it would be fitting, so the battery reports both and the pooled number for each.

Round 11, per-army median ρ with list share: 0.173 over 27 armies (round 11's report: 0.173). Five armies are negative (Space Wolves −0.18, Dark Angels −0.11, Drukhari −0.10, Chaos Daemons −0.08, Chaos Space Marines −0.04) and five more are under 0.10 (Aeldari, Orks, Death Guard, Blood Angels, Imperial Knights); the Tyranids (0.55) and Grey Knights (0.55) lead.

The ceiling: no pair yet. No two panels grade the same army, so there is no √ρ between reviewers to cap expectations. The code path is in place (ceilings()): when two panels name one faction, it reports their ρ on the units both grade and its square root.

4. The verdict: near zero on several armies

The plan asked one question of this battery (section 10, step 1): is round 11 in the right region (beats list share and points on most panels), in which case we fit its weights, or near zero on several armies, in which case we redesign the jobs? The answer is the second:

What this means for the queue: steps 3 and 4 (the matrix, the allocation rule, compose over all armies) come before the fitter, as planned, and the fitter's first job is to beat these numbers, not round 11's. The plan's success line stands as written: held-out ρ above list share's 0.21, or a residual we explain where list share explains nothing (the Space Marine panel, where list share is 0.07 and cheap points 0.34, is the first place to look).

5. Locked panels

A panel file with "locked": true at its top level is never scored by the battery: it prints the panel's unit count and date and nothing else, unless run with --final-check (the one look at the end, plan section 10 step 8). tests/unit/tierbaseline.test.js checks both. None of the three panels here is locked; the locked pair waits for the lists Jordan supplies.

6. Data and reproducibility