6 Oct 2026, Pacific (T-739). Frozen snapshot 8dcd24032d, inputs hash dcd477d8a78c (node tools/ledger-inputs.js --verify: 27 inputs match). Tools: tools/recipestep.js (the measurement), tools/reach.js (the Reach line), tools/postchain.js (post steps in a chain), tools/movecheck.js (the speed check). Numbers in build/recipe/ (git-ignored). Each prediction and keep rule was written in reports/tiers/ledger-steps.md before its scoring. Nothing on the page or in the published benchmark moves: adopting a kept step is Jordan's call.
In plain English
Step 1 (step 41), the stable layer: kept, as "Track A2 (candidate)". Track A plus the three dials the fold re-run picked in nearly every fold: charEpic 0.5 (an Epic Hero is worth a little more), aircraftBoard 1 (aircraft are off the board when objectives score) and typicalSize 0.25 (a unit ranks a quarter of the way toward its usual squad size).
- Held out on the 15 exam armies: +1.2 points [−0.4, +2.6], 11 of 15 armies not falling (9 up, 4 down). Leave one army out gives the same +1.2. With Gemini at half weight, +1.1.
- The exam's pooled score goes from 59.3 to 60.5. The Tyranids gain 1 point (82.6 on the three lists, consensus 90.4).
- On shuffled letters it does nothing (−0.2 to +0.3), so the gain is signal.
- It passes the keep rule written before scoring: a positive gain, an interval whose low end is above −0.5, and at least 9 of 15 armies not falling. It is weak, though: the interval touches 0.
- Winners: Blood Angels +8.1 (Astorath and Lemartes, graded A, climb), Agents +4.0, Drukhari +3.4, Thousand Sons +3.2 (Ahriman from 18th to 8th), Orks +2.6. The loser is Space Wolves, −5.7: their Epic Heroes, Ulrik and Ragnar, rise, but the reviewer grades them C.
- Saved as tools/ledger-benchmarks/2026-10-06-step41-stable.json. It is the base for step 2.
Before step 2: is the reviewers' "Move" really about speed? Movement had failed before (the every-dial screen: arrival reach +0.0, r.movement −0.2, charMove −0.24). So I checked first, on T-736's feature table (tools/movecheck.js; leave one army out; the same pair scoring).
- It is speed, but hidden behind unit type. On its own, Move barely orders the pairs (53.9%, ties as half).
- Fast units are more often vehicles and flyers and less often infantry. Reviewers mark the first two down and the third up. Once those types are held fixed, Move's weight triples and it adds +1.8 [+0.1, +3.5] over the ledger (10 armies up, 4 down).
- Per 100 points, it adds nothing over the ledger (−0.9 [−2.3, +0.0]). That is the scale every ledger line uses, and it explains why the old movement dials failed.
- So the check predicted step 2 as briefed would fail. It also suggested one variant: speed against units of the same kind, counted per unit. I declared that variant as step 42b, its own step with its own prediction, before scoring either one.
Step 2 (step 42), the Reach line as briefed: reverted. The fitted wReach is 0.
- The line is the share of the battlefield a unit can stand on by the end of its second turn. It takes the best of the unit's ways in: Move plus an advance, Scouts, Infiltrators, Deep Strike, reserves, a transport. It is priced per 100 points against the army's median.
- All 9 training folds, and all 15 leave-one-out fits, chose weight 0. Every weight loses when fixed: −0.4 at the smallest, −7.0 at the largest. Astra Militarum, Votann and Orks lose most.
- Shuffled letters do the opposite: they pick large weights and gain up to +3.6. The real letters reject what noise accepts.
Step 42b, Reach against its own kind, per unit: reverted.
- Held out: +0.6 [−1.7, +2.7], but only 8 of 15 armies don't fall (the rule needs 9). Gemini at half gives −0.3.
- Shuffled letters gain about as much (+0.2, +1.0, −1.7), so this is noise-sized.
- It does move the right units in some armies. The Land Speeder goes from 29th to 2nd in Space Marines and from 34th to 9th in Dark Angels (both graded S). Wolf Scouts go from 18th to 2nd.
- It moves the wrong ones in others: Thunderkyn and Hearthguard fall (both graded A), and World Eaters lose 6.6.
- The training folds chose wReach 4 (6 folds) and 2 (3 folds). That is the fitted weight, but the step is reverted, so nothing is saved.
What step 3 (durability) should look at first (the same check, run on durability facts; numbers below):
- Durability shows the opposite pattern to speed.
- Per 100 points, Toughness, toughness against our attackers, Wounds and the Soak line all add a little over the ledger on the 15 non-Tyranid armies: Toughness +2.6 [−0.8, +5.3], Soak +1.4, toughness against our attackers +1.7, Wounds +1.5. The Tyranids pull the other way, and every interval touches 0.
- Per unit, the same facts subtract (−0.8 to −1.2).
- The per-point gain mostly vanishes beside the type controls (−0.7 to +0.4), so part of it is unit type again.
- So step 3 should start with Soak or Toughness per 100 points, on the ledger's existing Soak line rather than a new one. It should run tools/movecheck.js
--featurefirst, and expect the Tyranids to resist. - The joint block "kind + speed within kind" is worth about +2.2 over the ledger in the check (61.9 → 64.1). That points to a kind layer (Infantry up, Fly and Monster down, Epic Hero up) as a candidate step of its own, not speed alone.
The test, as written before scoring (every step)
- Base: the previous kept step (Track A for step 41; step 41 for steps 42 and 42b, which are two definitions of one line).
- Measure: tools/recipestep.js. Each setting goes through the exam's own path (exam.js runOf, wCarry 0, its post).
- Human yardstick: each army's 11th-edition lists' mean pairwise accuracy, the exam's score; checked equal to exam.js, army by army.
- Held out: the 15 exam armies, 3 seeds × 3 folds of 5 (yardexam.js folds()). On each fold the best candidate is chosen on 10 armies and scored on the other 5. The base is always a candidate, and wins ties.
- Intervals: a paired unit bootstrap (1,000 draws; the picks as chosen), and armies resampled.
- Beside it: leave one army out; Gemini at half weight (exam --gemini 0.5); the Tyranids' 3 lists, 5 lists and their three-list consensus (583 pairs).
- Control: three shuffled letter sets (seeds 11 to 13). Each army's graded units are permuted jointly, each taking another's row of letters. The same choose-on-10, score-on-5 runs on them.
- Keep if the human held-out gain is above 0, its unit interval's low end is above −0.5, and at least 9 of 15 armies don't fall. Otherwise revert.
Step 41: the stable layer
| Yardstick | Held-out gain over Track A [units] {armies} | Up / down (not falling) | Leave one out | Folds' picks |
|---|---|---|---|---|
| Human lists | +1.2 ± 0.8 [−0.4, +2.6] {−0.3, +2.6} | 9 / 4 (11) | +1.2 [−0.4, +2.6] | step 41 ×9 |
| Gemini at half | +1.1 ± 0.8 [−0.4, +2.5] {−0.4, +2.5} | 9 / 5 (10) | +1.1 | step 41 ×9 |
| Shuffled, seeds 11 / 12 / 13 | +0.3 / −0.2 / −0.2 | +0.5 / +0.0 / +0.0 | mostly Track A |
| Exam pooled | Tyranids 3 lists | 5 lists | Tyranid consensus | |
|---|---|---|---|---|
| Track A | 59.3 | 81.6 | 78.8 | 89.5 |
| Step 41 | 60.5 | 82.6 | 79.4 | 90.4 |
Per army, held out (= fixed here; the folds always pick step 41): Space Marines +0.2, Adeptus Mechanicus −0.9, Agents +4.0, Astra Militarum +1.8, Blood Angels +8.1, Dark Angels +0.1, Death Guard +1.1, Drukhari +3.4, Emperor's Children +0.0, Genestealer Cults −0.2, Leagues of Votann +0.0, Orks +2.6, Space Wolves −5.7, Thousand Sons +3.2, World Eaters −0.1.
Prediction: +1.0 to +2.5, about 10 of 15 not falling, Tyranids within ±0.5, shuffles about 0. Right on the size, the count and the shuffles. The Tyranids gained +1.0, a little more than predicted. Space Wolves were predicted up and fell. One leak to disclose: the tool's check against exam.js printed step 41's per-army scores before the prediction was written. It showed no base scores and no changes.
The speed check (tools/movecheck.js)
These are pairwise logistic fits on T-736's table, leave one army out, with cascade2's consensus pairs and ties counted wrong. The "controls" are Infantry, Vehicle, Monster, Fly, points, Character and Epic Hero.
| Move's added held-out accuracy | 16 armies [armies resampled] | 15, no Tyranids | Up / down |
|---|---|---|---|
| over the controls | +2.0 [−0.8, +4.6] | +1.4 | 11 / 5 |
| over the ledger | +0.8 [−1.7, +3.4] | +0.8 | 10 / 6 |
| over the ledger and the controls | +1.8 [+0.1, +3.5] | +1.9 [+0.2, +3.7] | 10 / 4 |
| Move per 100 points, over the ledger | −0.9 [−2.3, +0.0] | −0.3 | 4 / 8 |
| Move per 100 points, over the ledger and the controls | +1.5 [−0.5, +3.6] | +1.5 | 9 / 6 |
- Move's coefficient: 0.17 alone, 0.51 beside the controls, 0.45 beside the ledger and the controls.
- Within an army, Move correlates with Vehicle (+0.41), Fly (+0.32), Toughness (+0.43) and Infantry (−0.55).
How the Reach line differs from the old dials.
- The old
reachswitch puts arrival reach on the Presence line, which is a few percent of Value. rules.movementmultiplies the Value of units that carry a movement rule.charMovecredits a character's movement rules.landsis the charge model's chance a melee unit gets there.- None of them credits a unit for being fast in itself. The Reach line does: the board it can stand on, from its Move, an advance and its ways in.
Step 42: Reach, as briefed
The measure (tools/reach.js; the ways, Move and transports are surefire.js's, refactored into basicsOf() so both read the same thing; sureFire's output is unchanged to the byte):
- r is the share of a 1" grid over the 44" × 60" board that the unit can be standing on by the end of our turn 2. The six deployment maps are weighted by their share of event layouts.
- Each way in covers this much:
- Walking: within 2(M + 3.5") of our zone.
- Scouts X": X more.
- Infiltrators: within 2(M + 3.5") of the ground more than 8" from the enemy's zone.
- Deep Strike and other arrivals: anywhere more than 8" from the enemy's zone (no move after; a turn-1 ingress moves once).
- Reserves: the 6" edge strip.
- A transport that fits the unit: the transport's two moves plus 3".
- Not counted: Fly (no terrain on the board), tunnels, advance-roll rules.
The line: Value + wReach × (r × 100 ÷ points − the army's median). wReach was chosen on each training fold from 0, 50, 100, 200, 400, 800 and 1,600.
| wReach (fixed) | Exam change [units] | Up / down | Half | Shuffled (mean of 3) | Tyranids 3 lists |
|---|---|---|---|---|---|
| 50 | −0.4 [−1.0, +0.1] | 3 / 8 | −0.4 | +0.2 | 82.4 |
| 100 | −0.5 [−1.4, +0.5] | 5 / 9 | −0.5 | +0.4 | 82.6 |
| 200 | −0.8 [−2.2, +0.6] | 7 / 8 | −0.5 | +0.5 | 82.6 |
| 400 | −2.3 [−4.6, −0.3] | 4 / 11 | −1.7 | +1.3 | 82.0 |
| 800 | −4.6 [−8.5, −0.9] | 4 / 11 | −3.8 | +2.0 | 79.2 |
| 1600 | −7.0 [−11.7, −2.5] | 4 / 11 | −5.8 | +2.0 | 75.1 |
- Held out: +0.0 (all 9 folds, and all 15 leave-one-out fits, pick 0).
- Shuffled letters held out: +3.6 / −2.0 / +0.7.
- Verdict: revert. The prediction (fit at 0, revert) was right.
Step 42b: Reach against its own kind, by the unit
The line: Value + wReach × P̃ × (r − the median r of the army's units of the same kind), where P̃ is the army's median Value. The kinds are Vehicle, Monster, Mounted, Infantry and other. The grid was 0, 0.25, 0.5, 1, 2, 4 and 8.
| Held out [units] {armies} | Up / down (not falling) | Leave one out | Picks | |
|---|---|---|---|---|
| Human lists | +0.6 ± 1.0 [−1.7, +2.7] {−1.3, +2.5} | 7 / 7 (8) | +0.6 [−1.9, +3.0] | 4 ×6, 2 ×3 |
| Gemini at half | −0.3 [−1.9, +1.2] | 8 / 7 (8) | −0.5 | 0 to 4 |
| Shuffled 11 / 12 / 13 | +0.2 / +1.0 / −1.7 | 1 to 8 |
- Fixed: 1 gives +0.1, 2 gives +0.6, 4 gives +0.9 [−1.7, +3.4], 8 gives −0.2. The Tyranids score 83.4 at 1 and 81.0 at 4.
- Per army, held out: Space Marines +2.9, Adeptus Mechanicus −4.7, Agents +5.7, Astra Militarum +2.6, Blood Angels +7.8, Dark Angels +4.5, Death Guard −3.2, Drukhari +2.9, Emperor's Children −0.7, Genestealer Cults −1.9, Votann −2.9, Orks +0.0, Space Wolves −0.2, Thousand Sons +2.3, World Eaters −6.6.
- Verdict: revert (8 of 15 not falling). The prediction was right on the size and the interval, but wrong on the count and on the shuffles (they were not about 0).
What the Reach line does to the units reviewers dispute
These are ranks within the army: step 41, then 42 at 200 (illustrative, since the fit chose 0), then 42b at 4 (the folds' pick). The letters are the 11th-edition lists'.
| Unit (army) | Letters | Step 41 | 42 at 200 | 42b at 4 | Reach r, Move |
|---|---|---|---|---|---|
| Defiler (DG) | B | 4 | 7 | 4 | 0.90, M 12 |
| Myphitic Blight-hauler (DG) | A | 11 | 11 | 12 | 0.82, M 10 |
| Foetid Bloat-drone (DG) | A | 16 | 13 | 13 | 0.82, M 10 |
| Plagueburst Crawler (DG) | C | 21 | 23 | 19 | 0.82, M 10 |
| Helbrute (DG) | C | 12 | 12 | 23 | 0.70, M 7 |
| Chaos Rhino (DG) | A | 33 | 35 | 29 | 0.90, M 12 |
| Brôkhyr Thunderkyn (Votann) | A | 5 | 5 | 7 | 0.88, M 5 (by Land Fortress) |
| Einhyr Hearthguard (Votann) | A | 16 | 17 | 17 | 0.88, M 5; Deep Strike only 0.57 |
| Sagitaur (Votann) | A | 15 | 12 | 11 | 0.96, M 12, Scouts |
| Wulfen (SW) | C | 8 | 6 | 6 | 0.78, M 9 |
| Wulfen with Storm Shields (SW) | B | 4 | 5 | 3 | 0.78, M 9 |
| Wolf Scouts (SW) | B | 18 | 17 | 2 | 1.00, Infiltrators |
| Land Speeder (SM) | S, S | 29 | 26 | 2 | 0.94, M 14 |
| Land Speeder (DA) | S | 34 | 33 | 9 | 0.94, M 14 |
| Scout Squad (SM) | S, S | 54 | 52 | 54 | 1.00 (Infiltrators), but most Marine infantry ride to 0.93–1.00 |
| Scout Squad (DA) | A | 53 | 51 | 47 | 1.00 (Infiltrators) |
- The Death Guard engines barely move under either line. Their reach (0.82 at Move 10) is their kind's median, so speed isn't what reviewers see in them. Only the slow Helbrute falls (graded C, so rightly). (T-737 found the Rhino's worth is delivery, Fire Support.)
- The kind variant fixes the Land Speeder and Wolf Scouts, but pushes Thunderkyn and Hearthguard the wrong way.
- Their reach comes from riding a Land Fortress, a way every Votann infantry unit shares.
- Deep Strike scores below walking on board share (0.57: it lands more than 8" from the enemy and can't move that turn).
- So the line can't see what reviewers value in an arrival: where it lands, not how much ground it covers.
- Scouts don't rise because transports lift every infantry unit to nearly full reach. A transport has room for one squad, but we count it for every rider.
Doubts and what I decided alone
- The horizon (turn 2) and the advance (+3.5" every move) are my choices, written before scoring. Turn 2 is where surefire.js measures. Counting every rider of a transport as using it flatters infantry, as it does in sureFire.
- Step 42b was my addition, prompted by the check. I declared it as its own step before any scoring. Two variants of one line were tried, so the multiplicity is two.
- The keep rule's −0.5 floor let step 41 through with an interval touching 0 (−0.4). The fold re-run had already chosen these dials, on cascade2's records rather than this measurement, so the evidence is consistent but not strong.
- Step 41's prediction leak is disclosed above.