Tactical Reroll
⋯

Step 1 closed: every panel, the baselines, two ceilings and the outcome signal (2026-10-03)

From reports/tiers/2026-10-03-step1.md , rendered when the site is built.

T-576 and the close of step 1 in design/tier-plan.md's queue (section 10). Command: node tools/tierbaseline.js --json reports/tiers/2026-10-03-step1.json (about a second; 2,000 bootstrap resamples, 2,000 shuffles, 200 random tie-breaks, fixed seeds). Inputs: docs/data/rankings.json, round 11's published JSON, reports/tiers/panel/. Round 11 is read, not re-run. Lists counted from grimstat-corpus (CC BY 4.0). A report only: nothing on the site changes.

In short. Five panels are now scored, on three armies. Pooled over the four panels no round was tuned on (166 units), list share alone gets 0.41 [0.28, 0.53] and round 11 gets 0.19 [0.04, 0.34]. Points alone get about 0 whichever way round they are read. Round 11 beats list share on one panel, the new Space Marine panel, and loses on the other four. On both armies that have two reviewers, the reviewers agree with each other far more than round 11 agrees with either: ceiling √ρ 0.83 for the Space Marines and 0.86 for the Tyranids. The verdict of T-570 stands: redesign the jobs (T-568, T-569) before any fit. The bar for the fitter is now list share's held-out 0.41, not 0.21.

1. The headline: pooled held-out ρ

"Held out" means no round was tuned on the panel. The headline pools Tista Minis' and Auspex's Space Marines, Sprues & Brews' Orks and Hivemind Hobbies' Tyranids, weighting each panel by its unit count; the interval resamples units within each panel. "All panels" adds Auspex's Tyranids, which steered rounds 5 to 11.

PredictorHeld-out pooled ρ [95%] (4 panels, 166 units)All five panels (210 units)T-570's held-out (2 panels, 75 units)
List share alone0.41 [0.28, 0.53]0.49 [0.38, 0.59]0.21 [−0.03, 0.44]
Points alone, cheaper is better0.02 [−0.13, 0.19]−0.01 [−0.16, 0.14]0.01
Points alone, dearer is better−0.02 [−0.19, 0.13]0.01 [−0.14, 0.16]−0.01
Frozen round 110.19 [0.04, 0.34]0.29 [0.16, 0.42]−0.03 [−0.27, 0.20]

Two caveats on the label. First, Hivemind's Tyranids grade the army we tuned on, through another reviewer. So round 11's 0.35 there is less of a surprise than its 0.41 on Auspex's Marines. Second, the fitting panels are held out only until the fitter runs. After that, the only truly unseen panel is the locked Astra Militarum list, and leave one army out is the honest score.

2. Every panel, with lift over its shuffle

Each cell gives ρ [95%] with its lift over that panel's shuffle in brackets, then within one on the reviewer's own letter shape, with its lift. The lift is the number minus the mean of 2,000 random orders of the same units.

Panel (units scored; shape S/A/B/C/D)List share alonePoints, cheaper betterPoints, dearer betterRound 11Shuffle: within one
Space Marines, Tista Minis (30; 4/11/9/6/0)0.07 [−0.30, 0.40] (+0.07); 22.0 (+0.4)0.34 [−0.06, 0.67] (+0.34); 27.2 (+5.6)−0.34 (−0.34); 17.5 (−4.1)0.14 [−0.22, 0.49] (+0.14); 24.0 (+2.4)21.6
Space Marines, Auspex (46; 11/17/11/5/2)0.34 [0.08, 0.56] (+0.34); 34.3 (+3.0)0.17 [−0.18, 0.48] (+0.17); 34.0 (+2.7)−0.17 (−0.17); 29.6 (−1.7)0.41 [0.12, 0.65] (+0.41); 38.0 (+6.7)31.3
Orks, Sprues & Brews (45; 3/13/25/3/1)0.30 [−0.03, 0.58] (+0.30); 39.1 (+1.4)−0.20 (−0.20); 36.0 (−1.7)0.20 [−0.10, 0.48] (+0.20); 40.0 (+2.3)−0.15 [−0.44, 0.17] (−0.15); 38.0 (+0.3)37.7
Tyranids, Hivemind Hobbies (45 of 49; 4/12/17/9/3)0.81 [0.68, 0.90] (+0.81); 44.5 (+13.3)−0.13 (−0.13); 30.5 (−0.7)0.13 [−0.20, 0.45] (+0.13); 33.3 (+2.1)0.35 [0.04, 0.62] (+0.35); 37.0 (+5.8)31.2
Tyranids, Auspex, tuned on (44; 13/7/13/4/7)0.81 [0.64, 0.91] (+0.82); 41.4 (+19.3)−0.11 (−0.11); 23.0 (+0.9)0.11 [−0.24, 0.45] (+0.12); 28.0 (+5.9)0.66 [0.42, 0.82] (+0.67); 36.0 (+13.9)22.1

3. Two ceilings

The ceiling is ρ between two reviewers on the units both grade, letters against letters, with ties averaged. √ρ is the most an engine can expect against either of them.

ArmyPanelsSharedρ [95%]√ρDates
Space MarinesTista Minis × Auspex150.69 [0.30, 0.91]0.8322 Jul and about 2 to 3 Sep: they straddle BSData's 30 Aug points update. Auspex's seven steps give 0.70 (panel file, T-575)
TyranidsHivemind × Auspex380.73 [0.53, 0.86]0.86About early July and about September, about two months apart. They straddle the points changes that dropped four units, and those four are left out here. All 41 shared units give 0.74 [0.55, 0.87], √ρ 0.86 (panel file)

Round 11 reaches 0.14 and 0.41 against the two Marine reviewers, where the ceiling is 0.83. On the Tyranids it reaches 0.35 against the untuned reviewer, where the ceiling is 0.86. List share alone, at 0.81, sits near the Tyranid ceiling, which suggests both Tyranid reviewers lean on what the lists take.

4. Round 11 against list share, per army

The median per-army ρ with list share is 0.173 over 27 armies. Five armies are negative and seven are within ±0.1 (unchanged; no new data).

5. The outcome signal (T-571, reports/tiers/2026-10-03-outcomes.md)

The outcome data covers 382 placed lists from 23 events across 33 armies, July to September 2026. It holds final placings only, with no game records. A unit's lift is the mean finishing percentile of its army's lists that take it, minus the mean of those that don't. Of 418 units with 3 or more lists on each side, 33 have a 95% interval clear of zero, where chance alone gives about 21. That is a real signal, but too weak to grade one unit. Its use is as an aggregate yardstick: the correlation of our letters with lift, weighted by the smaller side. No faction-versus-faction table may be read by a machine today. It is not yet compared with the panels or round 11 (next step).

6. Panel inventory

FileArmyReviewer, dateGradedUse
2026-10-02-space-marines-tistaminis.jsonSpace MarinesTista Minis, 22 Jul 202630fitting
2026-10-03-space-marines-auspex.jsonSpace MarinesAuspex Tactics, about 2 to 3 Sep 202646 (Company Champions unmatched)fitting
2026-10-02-orks-spruesandbrews.jsonOrksSprues & Brews, 22 Aug 2026 (codex review)45fitting
tyranids-auspex.jsonTyranidsAuspex Tactics, about Sep 202644fitting; tuned on (rounds 5 to 11)
2026-10-03-tyranids-hivemind.jsonTyranidsHivemind Hobbies, about early Jul 202649 (4 patch drops; Sky-slasher Swarms unmatched, Legends)fitting
2026-10-03-astra-militarum-auspex.jsonAstra MilitarumAuspex Tactics, about mid Sep 202661locked: counted, not scored, until the final check
2026-10-03-leagues-of-votann-discussion.jsonVotanna video discussion, about early Apr 2026 (10th edition)13reference: never scored, fitted or compared with the engine (locked on 3 Oct, unlocked once dated)
2026-10-03-leagues-of-votann-2025.jsonVotanna tier list video, about Oct 2025 (10th edition)17reference, as above

One locked panel, not two. The plan asked for two locked panels. 11th edition came out in June 2026, which makes both Votann lists 10th-edition, and Jordan has no more lists. So the second lock is waived (the supervisor's call; Jordan: "use what you can"). The final check rests on Astra Militarum alone, 61 units, which is wide enough for one ρ with an interval of about ±0.25.

Patch alignment. Points were read from BSData's own files at the reviewers' commits at run time. Nothing of theirs is kept in the repo, and the newest read matches rankings.json exactly. Hivemind: 2 Jul to 30 Sep checked, four units dropped. Auspex's Space Marines: 2 Sep to 30 Sep checked, no move, so T-575's gap is closed.

Addendum, 3 Oct: the panels re-sorted after the design review (T-580)

Command: node tools/tierbaseline.js --sanity --json reports/tiers/2026-10-03-battery-t580.json (same seeds, same data: BSData 374f505, corpus 570b8c6, before the 1 Oct dataslate). Jordan's answers to the design review, passed on by the supervisor, change four things:

Pooled

SetEntrants' list sharePoints, cheaper betterRound 11
Held out: two Space Marine panels (76 units)0.23 [0.00, 0.43]0.24 [−0.04, 0.46]0.30 [0.07, 0.50]
Same army, new reviewer: two Tyranid lists (86)0.72 [0.60, 0.81]−0.18 [−0.41, 0.06]0.38 [0.15, 0.56]
Unseen, both sets (162)0.49 [0.37, 0.59]0.01 [−0.16, 0.19]0.34 [0.18, 0.48]
Every fitting panel (206)0.56 [0.46, 0.65]−0.01 [−0.17, 0.14]0.41 [0.27, 0.53]

Per panel (ρ [95%] (lift over the shuffle))

Panel (class; units)Entrants' list sharePoints, cheaper betterRound 11 (SE)Round 11 within one: ours / shuffle
Space Marines, Tista Minis (held out; 30)0.07 [−0.30, 0.40] (+0.07)0.34 [−0.06, 0.67] (+0.34)0.14 [−0.22, 0.49] (+0.14), SE 0.1924 / 21.6
Space Marines, Auspex (held out; 46)0.34 [0.08, 0.56] (+0.34)0.17 [−0.18, 0.48] (+0.17)0.41 [0.12, 0.65] (+0.41), SE 0.1338 / 31.3
Tyranids, 2026-10-03-tyranids-hivemind.json (same army, new reviewer; 45)0.81 [0.68, 0.90] (+0.81)−0.13 [−0.45, 0.20] (−0.13)0.35 [0.04, 0.62] (+0.35), SE 0.1437 / 31.2
Tyranids, 2026-10-03-tyranids-second.json (same army, new reviewer; 41)0.62 [0.40, 0.78] (+0.62)−0.23 [−0.53, 0.11] (−0.24)0.41 [0.07, 0.68] (+0.41), SE 0.1434 / 27.3
Tyranids, Auspex (tuned on; 44)0.81 [0.64, 0.91] (+0.82)−0.12 [−0.45, 0.24] (−0.11)0.66 [0.42, 0.82] (+0.67), SE 0.1036 / 22.1
Orks, Sprues & Brews (sanity list, retired; 45)0.30 [−0.03, 0.58] (+0.30)−0.20 [−0.48, 0.10] (−0.20)−0.15 [−0.44, 0.17] (−0.15), SE 0.1538 / 37.7

Ceilings

ArmyPairSharedSame letterρ [95%]√ρ
Space MarinesTista Minis × Auspex1540.69 [0.30, 0.91]0.83
Tyranidsthe -hivemind file × Auspex38170.73 [0.53, 0.86]0.86
Tyranidsthe -second file × Auspex35100.67 [0.44, 0.82]0.82
Tyranidsthe -hivemind file × the -second file40140.54 [0.28, 0.72]0.74

The Tyranid ceiling, three reviewers: √ of the mean pairwise ρ, √0.65 ≈ 0.80. Before the patch drops the two new lists agree at 0.57 on 43 shared units (16 the same letter); the three dropped units take that to 0.54 on 40. The two new lists agree with each other less than either agrees with Auspex, and their tier names, shapes (4/12/19/11/3 against 6/11/16/9/2) and graded units differ: nothing in the numbers suggests one video summarised twice, which fits Jordan's ruling.

What it says. On the armies we never tuned on, round 11 (0.30) edges entrants' list share (0.23), inside the margin. On the Tyranids, list share is far ahead (0.72 against 0.38) and both new reviewers lean on it. Pooled over the unseen panels list share leads, 0.49 against 0.34. The bar for the fitter is list share's held-out ρ: 0.23 on the held-out set as now defined, 0.49 over everything unseen.