5 Oct 2026. tools/spanmatch.js and node tools/exam.js <settings> --span. A yardstick, not a step: no ledger value, setting or dial changes. Snapshot eb3590a (BSData 374f505), the ledger files as they stood before T-687's re-freeze.
In plain English
Jordan's rule: when a unit's squad sizes land on different letters (say S at 20 models and C at 10) and we don't know which size a reviewer graded, count it as matched if the reviewer's letter is one of ours. With the guard he agreed, a match earns 1 ÷ the number of letters our span covers (S to C covers four letters, so a quarter), so a sharp answer that matches earns more than a vague one.
What we found: the span changes almost nothing, because few units have one. Only about one graded unit in seven has squad sizes that land on different letters (81 of 597 on Track A, 101 on Track B, 120 on the trio). For those that do, the span wins some and loses about as many: a unit we put on the reviewers' letter at its best size loses half its credit when its other size lands a letter away. Pooled over the 14 armies, the span score equals the strict letter score to within half a point on all three settings. Both sit well above the placebo (spans of the same width placed at random), so the letters carry real signal, but the span adds nothing to it. The prediction was wrong: the span is not a few points above strict; Carnifexes are matched by their span (half credit, B at two models, C at one, every reviewer C), Hormagaunts are not (their sizes span B, A–C or B–C, and the reviewers say S or A).
The numbers
Each cell: span / strict letter / placebo, in % (mean credit per list; an army the mean of its lists; pooled as the exam pools, 14 armies, Leagues of Votann apart). Intervals: 95%, paired bootstraps on the exam's own draws (2,000, seed 1; the Tyranids on their own draws, seed 3, since the exam never draws them).
| Track A (step 34b) | Track B | Trio (A + aircraft, small pieces, characters attached) | |
|---|---|---|---|
| Tyranids (5 lists, 52 units) | 36.5 / 36.7 / 22.3 | 38.7 / 38.1 / 21.9 | 33.0 / 30.7 / 22.1 |
| Tyranids: span − strict | −0.2 [−4.5, +3.9] | +0.6 [−2.9, +4.4] | +2.3 [−0.8, +6.2] |
| Tyranids: span − placebo | +14.2 [+5.3, +23.3] | +16.8 [+8.9, +25.2] | +10.9 [+2.8, +19.4] |
| Pooled, 14 armies | 25.2 / 25.2 / 18.9 | 26.5 / 26.8 / 19.2 | 26.1 / 26.3 / 19.3 |
| Pooled: span − strict | −0.1 [−1.4, +1.1] | −0.3 [−1.8, +1.1] | −0.2 [−1.7, +1.3] |
| Pooled: span − placebo | +6.3 [+2.4, +10.2] | +7.3 [+3.5, +11.3] | +6.8 [+3.2, +10.8] |
| Δ span against A, pooled | – | +1.4 [−2.6, +5.3] | +0.9 [−2.7, +4.7] |
| Δ strict against A, pooled | – | +1.6 [−3.1, +6.1] | +1.0 [−2.9, +5.2] |
| Δ span against A, Tyranids | – | +2.2 [−5.3, +10.0] | −3.5 [−8.2, +0.5] |
| Leagues of Votann (apart) | 17.0 / 13.6 / 22.8 | 13.6 / 13.6 / 20.4 | 16.5 / 12.7 / 24.7 |
| Units with spans wider than one letter (all 16 armies) | 81 of 597 | 101 | 120 |
Per army (span / strict / placebo; Track A, B, trio):
| Army (lists, graded) | A | B | Trio |
|---|---|---|---|
| Space Marines (2, 84) | 23.2 / 23.4 / 18.0 | 26.0 / 25.7 / 17.9 | 27.4 / 27.1 / 18.4 |
| Adeptus Mechanicus (1, 29) | 50.0 / 48.3 / 21.3 | 22.4 / 10.3 / 23.6 | 31.6 / 24.1 / 24.5 |
| Agents of the Imperium (1, 26) | 19.2 / 19.2 / 16.7 | 38.5 / 38.5 / 17.3 | 32.7 / 34.6 / 16.0 |
| Astra Militarum (1, 61) | 17.2 / 18.0 / 18.3 | 28.7 / 31.1 / 17.8 | 31.1 / 36.1 / 18.3 |
| Blood Angels (1, 19) | 36.0 / 42.1 / 15.8 | 44.7 / 42.1 / 18.4 | 42.1 / 42.1 / 17.5 |
| Dark Angels (1, 77) | 22.7 / 23.4 / 21.3 | 26.0 / 22.1 / 21.4 | 27.3 / 26.0 / 20.4 |
| Death Guard (1, 36) | 20.4 / 19.4 / 24.1 | 20.8 / 22.2 / 22.7 | 15.7 / 16.7 / 23.6 |
| Drukhari (1, 23) | 14.5 / 13.0 / 22.5 | 18.8 / 21.7 / 20.8 | 12.3 / 8.7 / 22.5 |
| Emperor's Children (1, 23) | 39.1 / 39.1 / 15.9 | 34.8 / 43.5 / 15.2 | 23.9 / 26.1 / 15.2 |
| Genestealer Cults (2, 24) | 16.0 / 17.1 / 18.0 | 20.3 / 21.5 / 18.4 | 13.1 / 14.9 / 17.7 |
| Orks (1, 45) | 27.8 / 24.4 / 21.9 | 30.0 / 28.9 / 22.2 | 36.7 / 35.6 / 23.7 |
| Space Wolves (1, 20) | 17.5 / 15.0 / 15.8 | 11.7 / 15.0 / 15.8 | 23.3 / 25.0 / 16.7 |
| Thousand Sons (1, 28) | 19.6 / 21.4 / 14.3 | 15.5 / 17.9 / 15.7 | 17.3 / 17.9 / 14.9 |
| World Eaters (3, 30) | 28.9 / 29.4 / 20.7 | 32.9 / 35.2 / 22.0 | 30.6 / 32.9 / 20.3 |
The full table with each army's Δ intervals is what node tools/exam.js tools/ledger-benchmarks/2026-10-05-step34b-baseline.json tools/ledger-benchmarks/2026-10-05-trackB.json tools/ledger-benchmarks/2026-10-05-cross-army-trio.json --span --md prints.
Which Tyranid units the span moves
From node tools/spanmatch.js --bench <setting> --md (credit averaged over the five lists; "strict" the share of lists on our ranked letter).
- Gain (Track A): Carnifexes (C at one, B at two; all C: 0 → 0.50), Tyranid Warriors with Melee Bio-Weapons (A at three, B at six; all B: 0 → 0.50). Track B adds Ranged Warriors (+0.40), Zoanthropes, Barbgaunts (+0.20), Hive Guard (+0.10); the trio Von Ryan's Leapers and Hive Guard (+0.10).
- Lose (Track A): Biovores (S, A, B at one, two, three; all S: 1 → 0.33), Tyrant Guard, Genestealers, Termagants (−0.10 to −0.20). Track B: Pyrovores (−0.44), Neurogaunts (−0.40), Hormagaunts and Tyrant Guard (−0.20).
- Hormagaunts are not helped: B at both sizes on A (span B), C to A on B, B–C on the trio; the reviewers say S, A, B, A, S. The span can't reach a letter we put no size on.
What it means, and what it doesn't
- The span is no free lift. The strict letter score is the harder test and the span equals it on average: the partial-credit guard works as designed (wider spans pay less), and the units whose sizes straddle a cut are as likely to lose a strict hit as to gain a half one. A dial must hold on both yardsticks (decisions.md, 5 Oct): today they agree.
- Letter scores are low because the cuts are ours. Strict letters sit near 25% pooled (chance about 19 to 20%) against near 60% pairwise: our letters are quantiles of each army (10% S, 20% A, 30% B, 25% C, 15% D) and the reviewers' shares differ. The pairwise score doesn't care about that; letter scores do. This yardstick is for differences between settings on the same cuts, not for comparing with pairwise.
- Previews: a setting with a
post()preview is scored on its adjusted Values (strict from the adjusted Value, the span's sizes from the forms, which a preview doesn't revalue). Checked that it runs with tools/typical.js's preview and leaves the other settings' columns untouched; no dial is measured here.
How it is built
- Span: each legal squad size of a unit (its own solo forms, one member) valued at its best detachment and loadout, cut by the run's own thresholds cutS..cutC (the ones its ranked row is cut by; the army isn't re-ranked per size); the ranked row's letter is always included; the span is every letter from the best to the worst of these. One size (or no solo form) spans one letter.
- Credit: 1 ÷ width when the reviewer's letter is in the span, 0 outside. Strict: 1 on the ranked row's letter. Placebo: each unit's width placed uniformly among the 6 − width positions, three seeds (1, 2, 3, mixed with the army's name) averaged.
- exam.js
--span: a second table after the first; without it the output is byte-identical (checked on a full run and in tests/unit/spanmatch.test.js); its bootstrap rides on the exam's draws (boot()'s hook draws nothing). - Doubts: the span is contiguous (S at one size and C at another counts A and B as inside, at a quarter each), as the card read Jordan's rule; a "set of the sizes' letters only" reading would credit fewer letters at a higher rate. The sizes are the ones the ledger has solo forms for (most units two, a few three, Biovores and Pyrovores 1 to 3).