6 Oct 2026, Pacific (Jordan approved, 6 Oct). Frozen snapshot b3c416933e, inputs hash 1b88e955b9b3; the ledger is Track A2 (tools/ledger-benchmarks/2026-10-06-step41-stable.json, its post step included, wCarry 0). Tools: tools/statlearn.js (the feature table, now with --bench and --out) and tools/ensconfirm.js (the confirmation). Numbers in build/ensconfirm/confirm.json (git-ignored). The recipe, the three sets, the pass rule and a prediction were written in reports/tiers/ledger-steps.md (row ensemble-confirm) and committed before any confirmation score was computed. A report: nothing in the ledger, the sliders, the letters or the site moves, and nothing is adopted.
In plain English
The question. T-736 tried 34 models. One of them was declared before the run: average the ledger's rank with a rank from a small "datasheet model" (Move, keywords such as Infantry, Epic Hero and Lone Operative, Toughness, Wounds, support rules; no ledger lines at all). It beat the ledger by +3.9 [+0.4, +7.3] on 15 held-out armies. With 34 tries, one interval clearing 0 could be luck. So this card checks it once, cleanly, on letters the fit never saw.
How.
- I froze the recipe first, exactly as T-736 declared it, and wrote down the sets, the pass rule and my prediction.
- I re-fitted it on today's snapshot, with Track A2 as the ledger, both inside the ensemble and as the baseline to beat. A check: my code reproduces statlearn.js's own fit to the last digit.
- Then I scored it on three sets the fit had not trained on:
- (a) Gemini's letters for 15 armies, each army scored by a fit on the other 15. T-736 had already looked at these, so they are weaker evidence.
- (b) The new Orks list (Auspex Tactics, from T-733's search), never used in any fit.
- (c) Leave one list out. In the four armies with two or three human lists (Tyranids, World Eaters, Space Marines, Genestealer Cults), each list in turn was held back from the fit entirely, then scored. That is 10 lists.
- The pass rule, set beforehand: on the human lists of (b) and (c) together, the gain over the ledger must be above 0, with its 95% interval above −0.5. The gain must also be positive on at least 2 of the 3 sets.
Verdict: confirmed, by the rule written beforehand. The margin is smaller than T-736's.
| Set | Ensemble | Track A2 | Gain [95%] | Up / down | Passes? |
|---|---|---|---|---|---|
| (a) Gemini, 15 armies (weaker: already seen) | 62.7 | 57.5 | +5.3 [+1.8, +9.1] | 10 / 4 armies | yes |
| (b) The Orks candidate list, 967 pairs | 69.3 | 67.7 | +1.6 [−7.8, +9.8] | one list, up | no (too wide) |
| (c) Leave one list out, 10 lists | 70.7 | 66.6 | +4.1 [+1.3, +7.3] | 8 / 2 lists | yes |
| (b) and (c) together, 11 human lists | 70.6 | 66.7 | +3.9 [+1.3, +6.6] | 9 / 2 lists | yes |
What to make of it.
- The signal is real, but modest. Every set moves the same way, and the human lists clear the bar with room to spare.
- The Genestealer Cults carry a lot of (c): +12 on each of their lists. Without them, (c) is +2.1 [+0.1, +4.2] and the 9 remaining human lists +2.1 [+0.3, +3.9]. That still passes, at about half the size.
- Against Track A2, the plain leave-one-army-out gain is smaller than T-736 measured against Track A: +2.4 [−0.8, +5.6] on the 15 armies without Tyranids (8 up, 7 down), not +3.9. That is not a confirmation set, but it matters. Track A2 already took one datasheet signal into the ledger: the Epic Hero dial (charEpic 0.5). So part of the gap had already closed. What is left is about 2 to 4 points.
- (b) can't decide anything on its own. One list's interval runs from −7.8 to +9.8. Scored by the fit that leaves the Orks out entirely, it is −0.3. It counts towards the direction, not the size.
- The weakest armies for the ensemble stay weak. Leave one army out, it loses to Track A2 on Drukhari (−6.3), Adeptus Mechanicus (−5.6), Death Guard (−4.2), Dark Angels, Astra Militarum, Emperor's Children, Space Marines and the Tyranids (−1.5 to −2.2; the ledger was tuned on the Tyranids, so that comparison flatters the ledger). It gains most on the Genestealer Cults (+15.2), Votann (+11.4), Thousand Sons (+11.3) and Space Wolves (+10.0).
- The average is better than either half. The datasheet model alone does worse than the ensemble on the human lists (68.9 against 70.7 on (c)) but better on Gemini (64.0 against 62.7). The ledger and the datasheet model are wrong about different units, and averaging the two ranks keeps most of each one's strengths.
- What the fit weighs on Track A2. This is the same order as T-736's. Every sign holds in all 16 leave-one-out fits:
- up: Move (+0.44), Infantry (+0.43), Epic Hero (+0.33), Lone Operative (+0.29), Psyker (+0.22), Toughness (+0.18), a command-point rule, an army-rule enabler (+0.18 each), Wounds and Deep Strike (+0.11);
- down: Fly (−0.27), many kinds of support rule (−0.26), a debuff rule (−0.20), Transport (−0.16), Vehicle and long range (−0.11).
- A diagnostic, never a target: winners' list share. It comes from internal tournament data, so only the result is given here: the ensemble tracks it a little more closely than Track A2 does. The numbers stay in the git-ignored build/ensconfirm/confirm.json.
Doubts.
- (c) is not fully independent. The held-back list's army still trains the fit through its other lists, and these four armies were in T-736's own model choice.
- (a) had been looked at.
- (b) is one list, a Gemini-drafted transcription of a video. T-733 rates its reliability low-medium.
- The clean evidence is the 11 human lists, and 4 of them come from two armies that move a lot (Genestealer Cults and World Eaters).
- New human lists for the single-list armies (T-733's waiting drafts) would be the next honest test.
Prediction, checked.
- Right:
- the re-implementation equals statlearn.js's fit;
- the re-fit is +2 to +4 (+2.4);
- Gemini is +3 to +6 and passes (+5.3);
- the Orks candidate is positive but too wide to pass (+1.6);
- Genestealer Cults are +6 to +12 (+12.0);
- the Tyranids are flat to a little down (−0.7).
- Wrong:
- World Eaters were higher (+4.9, not +1 to +3), and Space Marines rose (+2.4, not −1 to −3);
- (c) was larger (+4.1, low end +1.3, not +1 to +3, low end about −1);
- the pooled human gain was larger (+3.9 [+1.3, +6.6], not +1 to +3, low end −1.5 to +0.5);
- so the verdict was confirmed, not the "mixed" I thought most likely.
- The share diagnostic went the way predicted, inside the predicted range.
Proposal: bringing the strongest signals into the ledger, explainably
The ensemble works, but "the average of two ranks" is not something a reader can follow unit by unit. The aim is to move its strongest signals into the ledger as named, priced pieces, one declared step at a time. Each step gets:
- one dial;
- base: the last kept step, starting from step 41 (Track A2);
- tools/recipestep.js's protocol: held out on the 15 exam armies, leave one army out, the unit bootstrap, Gemini at half, shuffled letters;
- its prediction and keep rule written before scoring.
One lesson comes first. Step 43, the unit-kind layer, put Infantry, Monster, Vehicle, Fly, Epic Hero, Psyker and speed into Value together on a 2,916-candidate grid, and it was reverted (+0.2). So no joint grids. Each signal gets its own step, with its own single dial, and is kept or dropped on its own.
The order, strongest and steadiest first:
- Speed within its kind. Move is the largest coefficient (+0.44, all 16 folds). T-739's move check showed it is speed masked by unit type, not a proxy. It is priced against the median Move of the army's units of the same body kind (Infantry, Mounted, Monster, Vehicle, other).
- Both reach steps (42, 42b) failed, so the form is the one the fit and step 43 agree on: Value × (M ÷ M̃ of its kind)^k.
- Step 43's folds held kSpeed steady at 1. Grid {0, 0.25, 0.5, 1}.
- Epic Hero and Psyker, one shared multiplier (+0.33 and +0.22, all 16 folds; step 43 picked +0.3 for each).
- Explainable as "a named character or a psyker carries rules and choices our dice don't see".
- It builds on charEpic 0.5, which is already in Track A2. The step is the extra beyond that, Psyker included.
- Grid {0, 0.15, 0.3}.
- Infantry (+0.43, all 16 folds; it splits when fitted next to the others). Explainable as "it can hide and stand on objectives where models of other kinds can't". One multiplier, on {0, 0.15, 0.3}.
- Toughness and Wounds, as durability. The fit gives T +0.18 and W +0.11, and T-736's lines-only fit gave Soak the most weight of any line (31%, against the ledger's 3%).
- The explainable form is the Soak line's weight (
wSoak), not new flags. It is one dial on a grid of ×1, ×2, ×4. - Only if that fails: Toughness against the army's median as a flag.
- The explainable form is the Soak line's weight (
- Lone Operative (+0.29, all 16 folds). Last, because the ledger already prices it (
rules.lone×1.93). The step only asks whether that multiplier should be larger, on {1.93, 2.5, 3}. Few units have the rule, so its interval will be wide.
Not proposed: Fly (−0.27) and the many support-rule kinds (−0.26). Both point down. They are worth a look as corrections of what the ledger already credits (aircraft, auras) only after the five above.
Card it? Not built here. If Jordan agrees, steps 1 to 5 can be one card, run one step at a time, each kept or reverted before the next starts.
The tables (tools/ensconfirm.js --md)
The sets against the ledger (Track A2)
Pairwise accuracy, ties wrong; gain = ensemble − Track A2 in points; 95% bootstrap intervals (2000 draws). Pass: gain above 0 and the interval's low end above −0.5.
| Set | Units scored | Ensemble | Track A2 | Gain [95%] | Up / down | Passes? | Datasheet model alone |
|---|---|---|---|---|---|---|---|
| (a) GEMINI AGGREGATE, leave one army out (bootstrap over armies; weaker: T-736 had looked) | 15 armies | 62.7 | 57.5 | +5.3 [+1.8, +9.1] | 10 / 4 armies | yes | 64.0 |
| (b) T-733 Orks candidate (Auspex Tactics), deployed fit (bootstrap over its units) | 1 list, 967 pairs | 69.3 | 67.7 | +1.6 [−7.8, +9.8] | 1 / 0 | no | 66.6 |
| (c) leave one list out, multi-list armies (bootstrap over the lists) | 10 lists | 70.7 | 66.6 | +4.1 [+1.3, +7.3] | 8 / 2 lists | yes | 68.9 |
| Pooled human: (b) and (c), each list once (bootstrap over the lists) | 11 lists | 70.6 | 66.7 | +3.9 [+1.3, +6.6] | 9 / 2 lists | yes | – |
Verdict by the frozen rule: confirmed. The pooled human set passes; the gain is positive on 3 of the 3 sets.
Sensitivity, not in the rule: (b) scored by the fit without Orks: 67.4 vs 67.7, −0.3 [−9.3, +8.1]. (c) bootstrapped over its 4 armies instead of its 10 lists: +4.1 [+0.7, +9.6]. The Orks candidate grades 54 ledger units (0 names unmatched, 0 without a letter); on the pairs it orders, the Orks' yardstick list agrees 76.8% of the time where it grades both differently.
(c) list by list
| Army | List | Pairs | Training pairs (its other lists) | λ | Ensemble | Track A2 | Gain [95%, its units] | Datasheet model alone |
|---|---|---|---|---|---|---|---|---|
| Tyranids | auspexFull | 1044 | 653 | 10 | 83.6 | 84.2 | −0.6 [−5.6, +4.7] | 75.5 |
| Tyranids | hivemindFull | 869 | 680 | 10 | 84.8 | 84.1 | +0.7 [−5.9, +6.3] | 79.4 |
| Tyranids | secondFull | 889 | 682 | 30 | 77.2 | 79.5 | −2.4 [−8.2, +3.1] | 71.5 |
| Space Marines | auspexFull | 2440 | 2375 | 10 | 63.9 | 62.4 | +1.5 [−6.1, +8.7] | 61.6 |
| Space Marines | tobias | 2375 | 2440 | 10 | 61.6 | 58.2 | +3.4 [−3.2, +10.1] | 60.5 |
| Genestealer Cults | rey | 207 | 177 | 30 | 58.9 | 46.9 | +12.1 [+0.0, +24.6] | 70.5 |
| Genestealer Cults | vonCarmian | 177 | 207 | 30 | 69.5 | 57.6 | +11.9 [−2.0, +28.3] | 73.4 |
| World Eaters | redPath | 284 | 260 | 10 | 67.3 | 60.2 | +7.0 [−6.3, +22.2] | 64.4 |
| World Eaters | tacticalSugar | 331 | 185 | 30 | 69.5 | 68.0 | +1.5 [−10.1, +12.6] | 62.2 |
| World Eaters | exalted | 350 | 211 | 10 | 70.9 | 64.9 | +6.0 [−5.2, +16.5] | 69.7 |
By army: Tyranids −0.7 (3 lists); Space Marines +2.4 (2 lists); Genestealer Cults +12.0 (2 lists); World Eaters +4.9 (3 lists).
The re-fit on Track A2, leave one army out (T-736's protocol; not a confirmation set)
16 armies: ensemble 65.4, Track A2 63.2, +2.2 [−0.9, +5.4]. 15 without Tyranids: 63.8 against 61.4, +2.4 [−0.8, +5.6], 8 up and 7 down. Re-implementation check against statlearn.js's own fit on the same table: largest difference per army 0.0e+0 points.
| Army | Consensus pairs | λ | Ensemble | Track A2 | Gain | Gemini pairs | Gemini: ensemble | Gemini: Track A2 |
|---|---|---|---|---|---|---|---|---|
| Tyranids | 583 | 3 | 88.9 | 90.4 | −1.5 | 629 | 69.3 | 72.0 |
| Space Marines | 1515 | 30 | 60.1 | 61.9 | −1.8 | 1253 | 60.0 | 57.9 |
| Adeptus Mechanicus | 322 | 3 | 64.9 | 70.5 | −5.6 | 173 | 75.1 | 65.9 |
| Agents of the Imperium | 223 | 10 | 68.6 | 66.4 | +2.2 | 167 | 70.1 | 62.9 |
| Astra Militarum | 1319 | 3 | 69.2 | 71.3 | −2.1 | 332 | 57.8 | 53.9 |
| Blood Angels | 124 | 3 | 75.0 | 69.4 | +5.6 | 51 | 58.8 | 39.2 |
| Dark Angels | 2325 | 10 | 51.4 | 53.7 | −2.2 | 185 | 69.7 | 60.0 |
| Death Guard | 476 | 3 | 53.6 | 57.8 | −4.2 | 187 | 58.3 | 56.1 |
| Drukhari | 207 | 10 | 57.5 | 63.8 | −6.3 | 159 | 52.2 | 56.6 |
| Emperor's Children | 191 | 10 | 67.5 | 69.6 | −2.1 | – | – | – |
| Genestealer Cults | 112 | 30 | 64.3 | 49.1 | +15.2 | 147 | 61.9 | 44.9 |
| Leagues of Votann | 105 | 10 | 72.4 | 61.0 | +11.4 | 50 | 46.0 | 46.0 |
| Orks | 606 | 10 | 69.0 | 66.2 | +2.8 | 399 | 61.9 | 65.7 |
| Space Wolves | 140 | 10 | 55.0 | 45.0 | +10.0 | 104 | 78.8 | 66.3 |
| Thousand Sons | 309 | 3 | 52.1 | 40.8 | +11.3 | 194 | 53.1 | 44.8 |
| World Eaters | 219 | 10 | 76.3 | 74.4 | +1.8 | 204 | 67.6 | 69.6 |
What the deployed fit weighs (all 16 armies, Track A2, λ 10)
Per army-SD of the feature on the logit scale; folds = how many of the 16 leave-one-army-out fits share the sign.
| Feature | Coefficient | Folds same sign |
|---|---|---|
| M | +0.44 | 16 |
| infantry | +0.43 | 16 |
| epic | +0.33 | 16 |
| lone | +0.29 | 16 |
| fly | −0.27 | 16 |
| supportKinds | −0.26 | 16 |
| psyker | +0.22 | 16 |
| debuffKind | −0.20 | 16 |
| T | +0.18 | 16 |
| cpKind | +0.18 | 16 |
| enabler | +0.18 | 16 |
| transportKw | −0.16 | 16 |
| W | +0.11 | 16 |
| vehicle | −0.11 | 16 |
| range | −0.11 | 16 |
| deepStrike | +0.11 | 16 |