Tactical Reroll
⋯

Confirming the datasheet model on lists none of the fits saw (T-748)

From reports/tiers/2026-10-06-ensemble-confirm.md , rendered when the site is built.

6 Oct 2026, Pacific (Jordan approved, 6 Oct). Frozen snapshot b3c416933e, inputs hash 1b88e955b9b3; the ledger is Track A2 (tools/ledger-benchmarks/2026-10-06-step41-stable.json, its post step included, wCarry 0). Tools: tools/statlearn.js (the feature table, now with --bench and --out) and tools/ensconfirm.js (the confirmation). Numbers in build/ensconfirm/confirm.json (git-ignored). The recipe, the three sets, the pass rule and a prediction were written in reports/tiers/ledger-steps.md (row ensemble-confirm) and committed before any confirmation score was computed. A report: nothing in the ledger, the sliders, the letters or the site moves, and nothing is adopted.

In plain English

The question. T-736 tried 34 models. One of them was declared before the run: average the ledger's rank with a rank from a small "datasheet model" (Move, keywords such as Infantry, Epic Hero and Lone Operative, Toughness, Wounds, support rules; no ledger lines at all). It beat the ledger by +3.9 [+0.4, +7.3] on 15 held-out armies. With 34 tries, one interval clearing 0 could be luck. So this card checks it once, cleanly, on letters the fit never saw.

How.

Verdict: confirmed, by the rule written beforehand. The margin is smaller than T-736's.

SetEnsembleTrack A2Gain [95%]Up / downPasses?
(a) Gemini, 15 armies (weaker: already seen)62.757.5+5.3 [+1.8, +9.1]10 / 4 armiesyes
(b) The Orks candidate list, 967 pairs69.367.7+1.6 [−7.8, +9.8]one list, upno (too wide)
(c) Leave one list out, 10 lists70.766.6+4.1 [+1.3, +7.3]8 / 2 listsyes
(b) and (c) together, 11 human lists70.666.7+3.9 [+1.3, +6.6]9 / 2 listsyes

What to make of it.

  1. The signal is real, but modest. Every set moves the same way, and the human lists clear the bar with room to spare.
  2. The Genestealer Cults carry a lot of (c): +12 on each of their lists. Without them, (c) is +2.1 [+0.1, +4.2] and the 9 remaining human lists +2.1 [+0.3, +3.9]. That still passes, at about half the size.
  3. Against Track A2, the plain leave-one-army-out gain is smaller than T-736 measured against Track A: +2.4 [−0.8, +5.6] on the 15 armies without Tyranids (8 up, 7 down), not +3.9. That is not a confirmation set, but it matters. Track A2 already took one datasheet signal into the ledger: the Epic Hero dial (charEpic 0.5). So part of the gap had already closed. What is left is about 2 to 4 points.
  4. (b) can't decide anything on its own. One list's interval runs from −7.8 to +9.8. Scored by the fit that leaves the Orks out entirely, it is −0.3. It counts towards the direction, not the size.
  5. The weakest armies for the ensemble stay weak. Leave one army out, it loses to Track A2 on Drukhari (−6.3), Adeptus Mechanicus (−5.6), Death Guard (−4.2), Dark Angels, Astra Militarum, Emperor's Children, Space Marines and the Tyranids (−1.5 to −2.2; the ledger was tuned on the Tyranids, so that comparison flatters the ledger). It gains most on the Genestealer Cults (+15.2), Votann (+11.4), Thousand Sons (+11.3) and Space Wolves (+10.0).
  6. The average is better than either half. The datasheet model alone does worse than the ensemble on the human lists (68.9 against 70.7 on (c)) but better on Gemini (64.0 against 62.7). The ledger and the datasheet model are wrong about different units, and averaging the two ranks keeps most of each one's strengths.
  7. What the fit weighs on Track A2. This is the same order as T-736's. Every sign holds in all 16 leave-one-out fits:
    • up: Move (+0.44), Infantry (+0.43), Epic Hero (+0.33), Lone Operative (+0.29), Psyker (+0.22), Toughness (+0.18), a command-point rule, an army-rule enabler (+0.18 each), Wounds and Deep Strike (+0.11);
    • down: Fly (−0.27), many kinds of support rule (−0.26), a debuff rule (−0.20), Transport (−0.16), Vehicle and long range (−0.11).
  8. A diagnostic, never a target: winners' list share. It comes from internal tournament data, so only the result is given here: the ensemble tracks it a little more closely than Track A2 does. The numbers stay in the git-ignored build/ensconfirm/confirm.json.

Doubts.

Prediction, checked.

Proposal: bringing the strongest signals into the ledger, explainably

The ensemble works, but "the average of two ranks" is not something a reader can follow unit by unit. The aim is to move its strongest signals into the ledger as named, priced pieces, one declared step at a time. Each step gets:

One lesson comes first. Step 43, the unit-kind layer, put Infantry, Monster, Vehicle, Fly, Epic Hero, Psyker and speed into Value together on a 2,916-candidate grid, and it was reverted (+0.2). So no joint grids. Each signal gets its own step, with its own single dial, and is kept or dropped on its own.

The order, strongest and steadiest first:

  1. Speed within its kind. Move is the largest coefficient (+0.44, all 16 folds). T-739's move check showed it is speed masked by unit type, not a proxy. It is priced against the median Move of the army's units of the same body kind (Infantry, Mounted, Monster, Vehicle, other).
    • Both reach steps (42, 42b) failed, so the form is the one the fit and step 43 agree on: Value × (M ÷ M̃ of its kind)^k.
    • Step 43's folds held kSpeed steady at 1. Grid {0, 0.25, 0.5, 1}.
  2. Epic Hero and Psyker, one shared multiplier (+0.33 and +0.22, all 16 folds; step 43 picked +0.3 for each).
    • Explainable as "a named character or a psyker carries rules and choices our dice don't see".
    • It builds on charEpic 0.5, which is already in Track A2. The step is the extra beyond that, Psyker included.
    • Grid {0, 0.15, 0.3}.
  3. Infantry (+0.43, all 16 folds; it splits when fitted next to the others). Explainable as "it can hide and stand on objectives where models of other kinds can't". One multiplier, on {0, 0.15, 0.3}.
  4. Toughness and Wounds, as durability. The fit gives T +0.18 and W +0.11, and T-736's lines-only fit gave Soak the most weight of any line (31%, against the ledger's 3%).
    • The explainable form is the Soak line's weight (wSoak), not new flags. It is one dial on a grid of ×1, ×2, ×4.
    • Only if that fails: Toughness against the army's median as a flag.
  5. Lone Operative (+0.29, all 16 folds). Last, because the ledger already prices it (rules.lone ×1.93). The step only asks whether that multiplier should be larger, on {1.93, 2.5, 3}. Few units have the rule, so its interval will be wide.

Not proposed: Fly (−0.27) and the many support-rule kinds (−0.26). Both point down. They are worth a look as corrections of what the ledger already credits (aircraft, auras) only after the five above.

Card it? Not built here. If Jordan agrees, steps 1 to 5 can be one card, run one step at a time, each kept or reverted before the next starts.

The tables (tools/ensconfirm.js --md)

The sets against the ledger (Track A2)

Pairwise accuracy, ties wrong; gain = ensemble − Track A2 in points; 95% bootstrap intervals (2000 draws). Pass: gain above 0 and the interval's low end above −0.5.

SetUnits scoredEnsembleTrack A2Gain [95%]Up / downPasses?Datasheet model alone
(a) GEMINI AGGREGATE, leave one army out (bootstrap over armies; weaker: T-736 had looked)15 armies62.757.5+5.3 [+1.8, +9.1]10 / 4 armiesyes64.0
(b) T-733 Orks candidate (Auspex Tactics), deployed fit (bootstrap over its units)1 list, 967 pairs69.367.7+1.6 [−7.8, +9.8]1 / 0no66.6
(c) leave one list out, multi-list armies (bootstrap over the lists)10 lists70.766.6+4.1 [+1.3, +7.3]8 / 2 listsyes68.9
Pooled human: (b) and (c), each list once (bootstrap over the lists)11 lists70.666.7+3.9 [+1.3, +6.6]9 / 2 listsyes–

Verdict by the frozen rule: confirmed. The pooled human set passes; the gain is positive on 3 of the 3 sets.

Sensitivity, not in the rule: (b) scored by the fit without Orks: 67.4 vs 67.7, −0.3 [−9.3, +8.1]. (c) bootstrapped over its 4 armies instead of its 10 lists: +4.1 [+0.7, +9.6]. The Orks candidate grades 54 ledger units (0 names unmatched, 0 without a letter); on the pairs it orders, the Orks' yardstick list agrees 76.8% of the time where it grades both differently.

(c) list by list

ArmyListPairsTraining pairs (its other lists)λEnsembleTrack A2Gain [95%, its units]Datasheet model alone
TyranidsauspexFull10446531083.684.2−0.6 [−5.6, +4.7]75.5
TyranidshivemindFull8696801084.884.1+0.7 [−5.9, +6.3]79.4
TyranidssecondFull8896823077.279.5−2.4 [−8.2, +3.1]71.5
Space MarinesauspexFull244023751063.962.4+1.5 [−6.1, +8.7]61.6
Space Marinestobias237524401061.658.2+3.4 [−3.2, +10.1]60.5
Genestealer Cultsrey2071773058.946.9+12.1 [+0.0, +24.6]70.5
Genestealer CultsvonCarmian1772073069.557.6+11.9 [−2.0, +28.3]73.4
World EatersredPath2842601067.360.2+7.0 [−6.3, +22.2]64.4
World EaterstacticalSugar3311853069.568.0+1.5 [−10.1, +12.6]62.2
World Eatersexalted3502111070.964.9+6.0 [−5.2, +16.5]69.7

By army: Tyranids −0.7 (3 lists); Space Marines +2.4 (2 lists); Genestealer Cults +12.0 (2 lists); World Eaters +4.9 (3 lists).

The re-fit on Track A2, leave one army out (T-736's protocol; not a confirmation set)

16 armies: ensemble 65.4, Track A2 63.2, +2.2 [−0.9, +5.4]. 15 without Tyranids: 63.8 against 61.4, +2.4 [−0.8, +5.6], 8 up and 7 down. Re-implementation check against statlearn.js's own fit on the same table: largest difference per army 0.0e+0 points.

ArmyConsensus pairsλEnsembleTrack A2GainGemini pairsGemini: ensembleGemini: Track A2
Tyranids583388.990.4−1.562969.372.0
Space Marines15153060.161.9−1.8125360.057.9
Adeptus Mechanicus322364.970.5−5.617375.165.9
Agents of the Imperium2231068.666.4+2.216770.162.9
Astra Militarum1319369.271.3−2.133257.853.9
Blood Angels124375.069.4+5.65158.839.2
Dark Angels23251051.453.7−2.218569.760.0
Death Guard476353.657.8−4.218758.356.1
Drukhari2071057.563.8−6.315952.256.6
Emperor's Children1911067.569.6−2.1–––
Genestealer Cults1123064.349.1+15.214761.944.9
Leagues of Votann1051072.461.0+11.45046.046.0
Orks6061069.066.2+2.839961.965.7
Space Wolves1401055.045.0+10.010478.866.3
Thousand Sons309352.140.8+11.319453.144.8
World Eaters2191076.374.4+1.820467.669.6

What the deployed fit weighs (all 16 armies, Track A2, λ 10)

Per army-SD of the feature on the logit scale; folds = how many of the 16 leave-one-army-out fits share the sign.

FeatureCoefficientFolds same sign
M+0.4416
infantry+0.4316
epic+0.3316
lone+0.2916
fly−0.2716
supportKinds−0.2616
psyker+0.2216
debuffKind−0.2016
T+0.1816
cpKind+0.1816
enabler+0.1816
transportKw−0.1616
W+0.1116
vehicle−0.1116
range−0.1116
deepStrike+0.1116