Tactical Reroll
⋯

The diagnosis pooled over all 16 armies, tested leaving one army out (T-635, 4 Oct)

From reports/tiers/2026-10-04-pooled-diagnosis.md , rendered when the site is built.

T-632's next step. T-632 found that on 52 Tyranid units every model lands at 68 to 73% on units it hasn't seen, and its learning curve was still rising. So this card asks whether more examples help: put every army's graded units through the same variables, and test on an army the model never saw. Command: node tools/diagnose.js --pooled (17.6 minutes on 4 worker threads; --quick for a 3-minute smoke run). Every number below is in 2026-10-04-pooled-diagnosis-tables.md and -pooled-diagnosis.json. This is a report: no ledger number, page default, data file, setting, benchmark or frozen input changed. Our own numbers; unit names only.

The short answer

The headline table

What's scored:

How it's measured:

ModelLeft out: pooledLeft out: army meanTyranidsSpace MarinesLeft out vs listsUnits held out, pooled (± SE)Scrambled, left out (SD)
Ceiling (each list against the others; 5 armies)––88.5%91.0%–––
(a) the ledger as it is (step 6)60.0%57.8%73.3%60.4%62.3%60.3% ± 0.249.1% (0.2)
(a1) the ledger's form, its 7 weights re-fitted59.9%57.8%73.2%60.4%62.2%60.1% ± 0.249.4% (0.2)
(b) linear on all 144 variables, ridge or lasso59.1%59.8%64.8%62.1%58.8%62.2% ± 0.550.9% (3.4)
(b2) linear on a short core list (25), ridge58.6%57.7%69.5%58.6%59.4%60.9% ± 0.448.1% (0.8)
(c) linear + every pairwise interaction (kernel)59.1%59.7%57.7%63.7%58.2%67.4% ± 0.650.1% (1.8)
(d) boosted trees, depth 1 to 3, monotone61.8%60.0%62.0%62.3%60.4%65.9% ± 0.749.3% (0.0)
(e) nearest neighbours (any army)63.4%\*56.3%57.9%69.6%\*61.0%65.4% ± 0.549.9% (2.6)

\* Twins: see the Space Marine family below.

Notes on the table:

Paired against the re-fitted ledger, leaving one army out (per-army gain in points; SE over the 16 armies):

PairMean gain ± SEArmies betterPooled gainScrambled spread (SD)
(d) trees − (a1)+2.2 ± 2.210 of 16+1.90.9
(b) linear − (a1)+2.0 ± 2.79 of 16−0.90.3
(c) interactions − (a1)+1.9 ± 3.010 of 16−0.90.6
(b2) core linear − (a1)−0.1 ± 2.16 of 16−1.42.6
(e) neighbours − (a1)−1.5 ± 2.87 of 16+3.4\*3.1

None is beyond one SE of 0. The flexible models win where the ledger is worst and lose where it is good. Below are the per-army wins and losses for (d), the trees, against (a), the ledger as it is:

Leave one army out, per army (against the consensus)

ArmyGradedCeiling(a)(a1)(b)(b2)(c)(d)(e)
Tyranids5288.5%73.3%73.2%64.8%69.5%57.7%62.0%57.9%
Space Marines8491.0%60.4%60.4%62.1%58.6%63.7%62.3%69.6%
Adeptus Mechanicus32–68.0%68.0%52.6%50.6%48.9%65.5%46.7%
Agents of the Imperium26–44.8%44.8%52.9%47.5%59.6%53.4%44.4%
Astra Militarum61–66.0%66.0%64.1%64.6%66.3%74.9%66.0%
Blood Angels19–54.0%54.0%71.8%53.2%67.7%54.0%43.5%
Dark Angels93–54.8%54.8%52.9%54.0%53.4%58.2%68.5%
Death Guard36–47.1%47.1%51.5%51.3%52.3%55.9%50.2%
Drukhari24–46.1%46.1%62.6%55.2%70.9%65.7%60.9%
Emperor's Children23–56.5%56.5%61.3%56.5%59.7%52.4%57.6%
Genestealer Cults2484.9%68.8%68.8%50.8%59.2%55.0%51.2%57.5%
Leagues of Votann2059.3%61.0%61.0%62.3%57.1%59.7%63.6%47.4%
Orks45–70.0%70.0%64.9%70.5%63.2%69.1%61.4%
Space Wolves21–46.9%46.9%60.6%66.9%56.3%54.4%58.1%
Thousand Sons28–45.6%45.6%61.2%53.1%58.6%54.0%55.3%
World Eaters3089.3%61.5%61.5%61.3%54.8%62.0%63.0%55.1%

The Space Marine family left out together

Why this check: the Dark Angels', Blood Angels' and Space Wolves' pools are their own datasheets plus the Space Marines' (tools/ledger-armies.js). Leaving one of the four out still trains on near-copies of most of its units. Here each of the four is scored with all four left out of the training (the main run's figure in brackets):

Army(a)(b)(b2)(c)(d)(e)
Space Marines60.4%55.7% (62.1)51.9% (58.6)56.6% (63.7)56.8% (62.3)55.5% (69.6)
Blood Angels54.0%59.7% (71.8)58.1% (53.2)62.9% (67.7)50.0% (54.0)58.1% (43.5)
Dark Angels54.8%46.3% (52.9)47.1% (54.0)47.1% (53.4)50.4% (58.2)52.4% (68.5)
Space Wolves46.9%62.5% (60.6)68.8% (66.9)56.3% (56.3)48.1% (54.4)56.3% (58.1)
The four, pooled57.1%51.1%49.9%51.8%53.2%54.0%

What it shows:

Units held out within armies (5 repeats × 5 folds, stratified by army)

All 16 armies are trained together on four-fifths of each army's units. The scores are on the pairs among the fifth left out of the same fit.

Army(a)(a1)(b)(b2)(c)(d)(e)
Tyranids73.2%72.9%67.8%68.7%68.0%66.9%62.9%
Space Marines61.1%60.9%61.5%58.3%69.6%65.9%70.4%
Adeptus Mechanicus70.1%69.8%52.3%56.3%52.1%64.7%55.1%
Agents of the Imperium43.3%43.4%51.1%51.9%58.9%44.2%47.5%
Astra Militarum66.5%66.4%71.5%74.8%78.9%77.5%74.9%
Blood Angels50.3%50.3%68.8%68.6%66.1%63.4%65.5%
Dark Angels54.9%54.8%59.0%56.7%66.1%65.2%67.0%
Death Guard46.8%46.6%52.8%52.2%58.0%57.5%54.5%
Drukhari49.1%49.1%63.1%49.7%70.8%69.2%58.1%
Emperor's Children53.1%52.3%60.5%54.4%66.5%56.1%58.3%
Genestealer Cults69.7%69.7%62.2%66.2%58.5%56.4%56.2%
Leagues of Votann65.4%65.4%61.3%56.6%63.6%57.8%52.0%
Orks67.9%67.7%68.0%69.9%66.1%69.4%57.6%
Space Wolves49.8%49.8%65.5%67.8%61.8%51.7%40.1%
Thousand Sons46.5%46.5%62.8%53.8%68.6%61.3%60.3%
World Eaters61.9%62.2%63.3%60.3%59.0%66.5%58.3%
Pooled60.3%60.1%62.2%60.9%67.4%65.9%65.4%

What it shows:

Does pooling help Tyranids?

Held-out pairwise on Tyranid units against their consensus, with the same folds in each column:

ModelTyranids alone (their own training folds)From the other 15 armies onlyThe other 15 plus Tyranids' own training folds
(a) the ledger as it is73.2%73.3%73.2%
(a1) re-fitted71.5%73.2%72.9%
(b) linear, all variables67.2%64.8%67.8%
(b2) core linear66.9%69.5%68.7%
(c) interactions65.4%57.7%68.0%
(d) trees66.0%62.0%66.9%
(e) neighbours61.6%57.9%62.9%

What it shows:

The learning curve over the number of training armies

Setup:

Model3691215
(d) trees59.5% / 58.1%61.4% / 58.8%55.1% / 56.5%65.5% / 59.4%61.8% / 60.0%
(e) neighbours57.1% / 54.1%59.8% / 55.9%56.5% / 55.3%65.5% / 58.2%63.4% / 56.3%
(b2) core linear52.4% / 53.3%59.7% / 58.3%56.6% / 57.3%58.0% / 57.1%58.6% / 57.7%
(a1) re-fitted ledger59.9% / 57.8%59.9% / 57.8%59.9% / 57.8%59.9% / 57.8%59.9% / 57.8%
(a) the ledger60.0% / 57.8%(no fit)60.0% / 57.8%

What it shows:

The why

What the pooled models key on (permutation importance on the army left out)

How it's measured: each army is scored by the fit on the other 15. A variable, or a group of variables, is shuffled among that army's units. The drop is in pooled pairwise, in points (each army weighted by its pairs). "Armies" counts how many armies drop by more than half a point.

FamilyTrees (d): droparmiesNeighbours (e): droparmies
Ledger lines4.8132.110
Weapons1.783.59
Kill by class1.4100.67
Datasheet1.3102.411
Rule flags0.7135.110
Interactions0.380.67

Reading the two models:

T-632's candidate mechanics, checked across armies

Each candidate is its group of variables (listed in tools/diagnose.js's CANDIDATES), shuffled together on the army left out. Beside that is each variable alone: the share of within-army pairs it orders right, and the armies where its sign is + or − (Spearman beyond ±0.1).

#CandidateTrees: drop (armies)Neighbours: drop (armies)Alone, the best of its variablesVerdict
2Scarce answers (Kill into monsters and vehicles, anti-tank guns)2.0 (7)2.2 (10)Kill into light vehicles 56.3% (+8/−2); ranged anti-vehicle 55.5% (+7/−6); monsters 54.3% (+6/−3)Backed: the strongest across armies
3Objective bodies (Score, OC per point, Battleline)0.8 (8)0.5 (8)bodies × OC × Move 56.7% (+9/−5); OC 56.0% (+8/−4)Backed, weakly. The trees use it as a floor: the bottom of an army's Score loses about half a letter
1Threat on arrival0.5 (7)0.5 (5)arrives × melee Kill 54.3% (+9/−2); arrives 54.1% (+9/−2)Weak. It points the right way in 9 armies but carries little. T-633's arrival Kill was reverted
8Value at the cheapest size−0.2 (2)1.0 (9)59.6% (+10/−1)Mostly the same signal as Value itself. Strong alone, nothing extra once Value is in
7Cheap reach (actions, indirect fire)0.2 (6)0.4 (10)points at the cheapest size 55.7% (+10/−1); actions line 47.2% (+4/−7)Not backed as actions; the actions line points down in 7 armies
4Character sniping (Precision)0.0 (3)0.2 (6)Kill into characters 54.6%; Precision 49.7%Not backed across armies (a Tyranid signal)
6Psychic0.1 (2)−0.1 (4)psyker 52.1% (+7/−3)Not backed. Psychic weapons are among the over-rated units' traits across armies
5Transports−0.2 (0)0.2 (6)transport 48.8% (+1/−6)Negative, as in T-632: the ledger over-credits transports' own stats. A transport's carrying is what's missing
9Melee Soak0.0 (0)0.1 (6)49.0% (+2/−8)Negative in 8 armies. The T-632 sign holds: melee staying power is over-credited

Ranked for the next ledger steps:

  1. Scarce answers (2).
  2. A capped Score (3).
  3. Arrival (1) only as part of a bigger change.
  4. Transports (5) as their carrying, and melee Soak (9) counted down. Both point the same way in most armies, but as corrections, not as new value.

Precision, psychic and cheap reach were Tyranid-specific.

The steadiest single signals (alone, within army)

VariablePairs rightArmies + / −
The Kill line (c_kill)60.1%11 / 0
Value at step 660.0%10 / 2
Value at the cheapest size59.6%10 / 1
Kill into heavy infantry58.8%10 / 0
Kill into cavalry and beasts58.6%11 / 1
√(Kill × Soak)58.1%8 / 4
The Hold line57.1%8 / 3

What it shows:

The trees on all 16 armies (depth 3, 100 trees; thresholds in SDs within the army)

The splits that carry the most gain:

The interactions (one split under the other):

So it is the ledger's Kill read differently by role, the same shape tools/roles.js tried with its fixed roles.

The units every flexible model misses the same way, leaving their army out

Each model's scores are mapped onto the army's letters by rank. The figure is the mean over the five flexible models minus the consensus, and all five miss in the same direction. The ledger's error is beside it.

Rated higher by the reviewersArmyConsensusModelsLedgerRated lower by the reviewersArmyConsensusModelsLedger
Stormraven GunshipDark AngelsA−4.0−4Centurion Assault SquadDark AngelsF+3.6+4
Miasmic MalignifierDeath GuardS−4.0−4Tzaangor EnlightenedThousand SonsD+3.4+4
BiovoresTyranidsS−3.9−0.6Serberys SulphurhoundsAdeptus MechanicusF+3.0+4
Suppressor SquadDark AngelsS−3.8−3Company HeroesDark AngelsF+3.0+3
Librarian in Phobos ArmourDark AngelsS−3.6−4LazarusDark AngelsF+3.0+3
LibrarianDark AngelsS−3.6−5Deathwing Terminator SquadDark AngelsD+3.00
SorcererThousand SonsS−3.6−4MaulerfiendThousand SonsC+3.0+3
JakhalsWorld EatersS−3.5−2.3Heavy Intercessor SquadDark AngelsD+2.8+3
VenomDrukhariS−3.4−5Grey Knights Terminator SquadAgents of the ImperiumC+2.6+3
Infernal MasterThousand SonsS−3.4−4Primaris PsykerAstra MilitarumC+2.6+1

What the misses share:

What this means

  1. Data is not the bind across armies; the reviewers' differences are. T-632 read its rising Tyranid curve as "more data would help". Pooled, the extra data comes from other armies with other reviewers and other tastes, and it doesn't help a new army. The ceiling the ledger can reach on an army it was never tuned on looks like about 60% pooled: what Kill and Value give.
  2. The ledger is not the problem on held-out armies. Its form, re-fitted on 15 armies, comes back as step 6. Every flexible model is within noise of it, and below it once the Marine twins are removed.
  3. The next measurable gains are the shared misses, not a better fit:
    • a leader's and a transport's worth to another unit (the joined-character and carrying measurements);
    • scarce anti-elite and anti-tank answers (T-632's candidate 2);
    • a Score floor rather than a slope.

Each would be one step, judged on lists and on units held out (--on).

  1. Within-army learning is real but is fitting a reviewer. A model trained on half an army's letters orders the other half 7 points better than the ledger in the armies the ledger reads worst. That is a per-army calibration of one reviewer's taste, not a rule of the game. It shouldn't move the ledger.

Method notes