Tactical Reroll
⋯

Headlines on 11th-edition lists (T-691)

From reports/tiers/2026-10-05-headline11e.md , rendered when the site is built.

5 Oct 2026. A yardstick step: no ledger value moves, only which reviewer lists the headline numbers average. Snapshot 8eb69bc7af, inputs hash a20dbd2c3316. Our own numbers; unit and reviewer names only.

What changed (Jordan, 5 Oct): the headlines score 11th-edition lists only, with 10th-edition lists shown beside them. 11th edition came out in June 2026. A new helper, tools/edition.js, decides which edition each panel in reports/tiers/panel/ belongs to. It reads the panel's edition field first, then its date (June 2026 or later is 11e), then its use line, then its source. Every panel gets an answer (node tools/edition.js lists them):

Measure (pairwise %, wCarry 0)Track ATrack BB + four dialsCross-army trio
Tyranid headline, 3 11e lists (old: 5 lists)82.4 (79.1)79.6 (77.7)83.8 (81.4)79.7 (76.5)
Held-out Tyranid units, diagnose (a), 10 × 5, vs the consensus (old: five-list consensus)81.2 (78.5)77.3 (77.2)82.4 (81.6)78.0 (75.7)
Headline lists, 3 Tyranid + 2 Space Marine, all 11e (old: the seven)73.6, ρ 0.531 (73.8, 0.545)71.3, 0.487 (72.3, 0.520)73.7, 0.528 (74.9, 0.564)71.8, 0.499 (71.8, 0.516)
Exam, 15 armies on their 11e lists (old: 14 armies)59.1 (59.0)60.0 (59.0)59.7 (59.0)61.9 (61.0)
Δ against A, units [95%]–+0.9 [−1.9, +3.6] (+0.0 [−2.6, +2.7])+0.5 [−2.3, +3.5] (+0.0 [−2.9, +2.9])+2.8 [+0.4, +5.4] (+2.0 [−0.4, +4.3])
Δ against A, armies resampled–+0.9 [−2.5, +4.1]+0.5 [−2.9, +4.0]+2.8 [+0.0, +5.6]
Armies better than A–7 of 15 (6 of 14)7 of 15 (6 of 14)9 of 15 (8 of 14)
Leagues of Votann, 11e Auspex list (18 units)61.073.368.675.2
Leagues of Votann, 10e lists beside (3 lists; not pooled; old row: all 4 lists, set apart)52.5 (54.6)63.1 (65.6)61.2 (63.0)63.6 (66.5)
Reviewers' ceiling on the Tyranid headline (each list against the others' mean letter)92.7 (90.4 on five)

How to read it:

A caution on Votann's Auspex list: it leans on popularity. It is not a lettered tier list. It groups units into three bands, standout, solid and niche, partly by how much they are played ("most played", "less played"), not only by how good they are. Popularity tends to favour units that are good in the current lists, which is part of what the reviewers measure, but it isn't the same thing. Its interval is wide ([−2.1, +33.7] for the trio).

My suggestion: treat the trio's move from "touches 0" to "clears 0" as not yet earned. Check it with the pool computed without Votann, which the exam prints as its own row: +2.0 [−0.4, +4.3], unchanged.

Run: node tools/exam.js tools/ledger-benchmarks/{2026-10-05-step34b-baseline,2026-10-05-trackB,2026-10-05-trackB-dials4,2026-10-05-cross-army-trio}.json --md (2,000 draws, seed 1); node tools/ledgerstep.js --base <setting>; tools/diagnose.js heldOut() with model (a), 10 repeats.