5 Oct 2026. A yardstick step: no ledger value moves, only which reviewer lists the headline numbers average. Snapshot 8eb69bc7af, inputs hash a20dbd2c3316. Our own numbers; unit and reviewer names only.
What changed (Jordan, 5 Oct): the headlines score 11th-edition lists only, with 10th-edition lists shown beside them. 11th edition came out in June 2026. A new helper, tools/edition.js, decides which edition each panel in reports/tiers/panel/ belongs to. It reads the panel's edition field first, then its date (June 2026 or later is 11e), then its use line, then its source. Every panel gets an answer (node tools/edition.js lists them):
- Tyranids: Auspex (full), HivemindHobbies (full) and Into the Hive Mind (full) are 11e. Maelstrom (March 2026) and Astrategas (April 2026) are 10e. The Tyranid headline is now the three 11e lists, with the five-list number beside it in
tools/ledgerstep.js,tools/diagnose.js(its consensus is now over the three;--fiverestores the old one) and the exam's Tyranid rows. - Leagues of Votann: Jordan's Auspex list (29 Aug) is 11e. Astrategas (May), the 2025 list and the April discussion are 10e. Votann rejoin the exam's pool on the Auspex list alone, and their 10e lists get a row of their own beside it.
- Every other army had only 11e lists, so its exam row is unchanged to the bit. So is the 14-army pool, still printed as "the T-680 pool".
- T-680's APART set is gone. An army now joins the pool when it has an 11e list. An army with none would be shown beside the pool automatically.
--all-editions(or the old--with-votann) scores every list together and reproduces the old 15-army numbers to the bit.
| Measure (pairwise %, wCarry 0) | Track A | Track B | B + four dials | Cross-army trio |
|---|---|---|---|---|
| Tyranid headline, 3 11e lists (old: 5 lists) | 82.4 (79.1) | 79.6 (77.7) | 83.8 (81.4) | 79.7 (76.5) |
| Held-out Tyranid units, diagnose (a), 10 × 5, vs the consensus (old: five-list consensus) | 81.2 (78.5) | 77.3 (77.2) | 82.4 (81.6) | 78.0 (75.7) |
| Headline lists, 3 Tyranid + 2 Space Marine, all 11e (old: the seven) | 73.6, ρ 0.531 (73.8, 0.545) | 71.3, 0.487 (72.3, 0.520) | 73.7, 0.528 (74.9, 0.564) | 71.8, 0.499 (71.8, 0.516) |
| Exam, 15 armies on their 11e lists (old: 14 armies) | 59.1 (59.0) | 60.0 (59.0) | 59.7 (59.0) | 61.9 (61.0) |
| Δ against A, units [95%] | – | +0.9 [−1.9, +3.6] (+0.0 [−2.6, +2.7]) | +0.5 [−2.3, +3.5] (+0.0 [−2.9, +2.9]) | +2.8 [+0.4, +5.4] (+2.0 [−0.4, +4.3]) |
| Δ against A, armies resampled | – | +0.9 [−2.5, +4.1] | +0.5 [−2.9, +4.0] | +2.8 [+0.0, +5.6] |
| Armies better than A | – | 7 of 15 (6 of 14) | 7 of 15 (6 of 14) | 9 of 15 (8 of 14) |
| Leagues of Votann, 11e Auspex list (18 units) | 61.0 | 73.3 | 68.6 | 75.2 |
| Leagues of Votann, 10e lists beside (3 lists; not pooled; old row: all 4 lists, set apart) | 52.5 (54.6) | 63.1 (65.6) | 61.2 (63.0) | 63.6 (66.5) |
| Reviewers' ceiling on the Tyranid headline (each list against the others' mean letter) | 92.7 (90.4 on five) |
How to read it:
- Tyranids: every setting rises 2 to 3 points on the three 11e lists. The 10e lists, Astrategas above all, were the ones we matched worst. The order of the settings doesn't change (B + four dials, then A, then the trio and B).
- The ceiling also rises, from 90.4 to 92.7. So the gap to the ceiling hardly closes: 11.3 points before, 10.3 now (Track A).
- The exam: adding Votann lifts the pool by about 0.1 point on A and about 1 point on the others. That is because Votann's one 11e list prefers B and the trio by 12 to 14 points.
- The trio's interval now clears 0 on units (+0.4 at the low end) and touches 0 with the armies resampled. That change rests on one army's one list of 18 units.
A caution on Votann's Auspex list: it leans on popularity. It is not a lettered tier list. It groups units into three bands, standout, solid and niche, partly by how much they are played ("most played", "less played"), not only by how good they are. Popularity tends to favour units that are good in the current lists, which is part of what the reviewers measure, but it isn't the same thing. Its interval is wide ([−2.1, +33.7] for the trio).
My suggestion: treat the trio's move from "touches 0" to "clears 0" as not yet earned. Check it with the pool computed without Votann, which the exam prints as its own row: +2.0 [−0.4, +4.3], unchanged.
Run: node tools/exam.js tools/ledger-benchmarks/{2026-10-05-step34b-baseline,2026-10-05-trackB,2026-10-05-trackB-dials4,2026-10-05-cross-army-trio}.json --md (2,000 draws, seed 1); node tools/ledgerstep.js --base <setting>; tools/diagnose.js heldOut() with model (a), 10 repeats.