Tactical Reroll
⋯

T-632 tables (tools/diagnose.js output)

From reports/tiers/2026-10-04-variables-headline-diagnosis-tables.md , rendered when the site is built.

Ceiling (each list against the mean of the other four): 90.4%

ModelIn-sample vs consensus (tuned)In-sample, most flexibleIn-sample vs listsHeld out vs consensus (± SE)Held out vs lists (± SE)Held out, pooled across folds (consensus / lists)Scrambled vs consensus (SD)Scrambled vs lists (SD)
(a) the ledger as it is (step 6)75.3%75.3%75.4%74.5% ± 0.675.0% ± 0.672.2% / 72.4%51.1% (4.3)51.4% (4.5)
(a1) the ledger's form, its seven weights re-fitted75.3%75.9%75.4%71.8% ± 1.171.9% ± 1.269.7% / 69.6%50.2% (4.5)50.5% (4.7)
(b) linear on every variable (ridge or lasso)96.9%99.6%92.5%70.4% ± 1.169.5% ± 0.969.3% / 68.5%45.2% (7.2)45.7% (7.2)
(b2) linear on a short core list (25 variables), ridge74.4%81.4%74.0%67.4% ± 0.965.7% ± 0.864.6% / 63.3%49.4% (7.5)49.7% (8.1)
(c) linear plus every pairwise interaction (kernel, ridge)95.0%97.6%91.7%72.4% ± 0.971.4% ± 0.872.1% / 71.2%45.6% (7.6)45.2% (8.2)
(d) boosted trees, monotone where the sign is obvious93.7%100.0%91.3%69.2% ± 0.968.4% ± 0.867.8% / 67.8%51.4% (5.3)51.5% (4.0)
(e) nearest neighbours (similar units)100.0%100.0%93.3%66.2% ± 1.764.9% ± 1.662.5% / 61.6%43.4% (8.8)43.0% (8.5)

Learning curve (trained on m units, scored on the pairs among the rest; ± SE over splits):

Per list, held out:

ModelauspexFullhivemindFullsecondFullmaelstromastrategas
ceiling92.1%92.9%87.7%95.8%83.7%
a80.9%80.1%78.3%70.8%65.1%
a178.3%76.8%75.1%67.2%62.0%
b74.7%72.2%72.7%65.5%62.6%
b272.1%70.7%66.4%58.7%60.7%
c76.6%74.1%74.7%68.1%63.7%
d74.5%73.0%73.2%62.0%59.5%
e70.2%68.1%69.7%59.3%57.4%

Settings chosen in the folds (most often):

Flexible against the ledger, held out (paired; SE over repeats; bootstrap over units SD and the share of resamples at or below 0):

Pairvs consensusvs listsbootstrap SDP(≤ 0)scrambled SD
a1 − a-2.6 ± 0.8-3.1 ± 0.91.71.002.1
b − a-4.0 ± 1.0-5.5 ± 0.74.60.789.7
b2 − a-7.1 ± 1.2-9.3 ± 1.25.50.899.2
c − a-2.0 ± 0.7-3.6 ± 0.74.60.679.9
d − a-5.2 ± 0.8-6.6 ± 0.74.00.917.6
e − a-8.3 ± 1.7-10.1 ± 1.74.70.9411.4
b − a1-1.4 ± 0.8-2.3 ± 0.83.90.669.0
b2 − a1-4.5 ± 1.4-6.2 ± 1.64.80.838.6
c − a10.6 ± 0.8-0.5 ± 1.13.80.509.5
d − a1-2.6 ± 1.1-3.5 ± 1.43.20.806.8
e − a1-5.7 ± 1.5-7.0 ± 1.54.10.9111.2

Permutation importance, held out, model c (drop in pairwise vs consensus, points):

VariableDropSE
kw_is_burrowers0.570.12
fx_arrival0.550.14
inv0.540.21
kw_precision0.440.16
kw_is_fly0.410.20
kw_is_titanic0.400.08
kw_heavy0.380.13
kw_is_transport0.370.12
rulesWeaponLike0.310.10
kw_is_battleline0.310.16
ix_invxT0.270.16
bestWS0.210.13
has_leader0.200.08
value0.200.19
fx_attMods0.200.05

The T-636 additions alone (top 12):

VariableFamilyDropSE
rulesWeaponLikerules0.310.10
bestWSloadout0.210.13
lfx_rulesrules0.180.13
scouts_inchesrules0.180.06
anti_infantryweaponkw0.170.14
transportCapacityrules0.170.12
dsMultiProfileloadout0.160.07
modelProfilessize0.140.09
rulesCountrules0.130.09
kwu_Transportkeywords0.130.15
maxAP_anyloadout0.110.13
ppmDiscountsize0.100.15
FamilyDropSE
ledger3.150.59
flags3.070.72
all T-636 additions1.860.57
weapons1.090.39
rules1.020.31
killclass0.950.31
keywords0.360.22
datasheet0.210.23
weaponkw0.030.23
size-0.150.22
interaction-0.370.23
loadout-0.760.36

Permutation importance, held out, model d (drop in pairwise vs consensus, points):

VariableDropSE
c_score3.600.54
value3.190.50
valueCheap1.240.33
k6_light_vehicle1.230.32
lands0.760.26
k6_monster0.480.17
inv0.390.16
abilitiesOwn0.340.11
c_hold0.310.13
kw_precision0.190.15
OC0.170.11
ppmDiscount0.160.16
ix_arrivexMelee0.130.26
carriedDistinct0.110.07
rangedVeh1000.100.10

The T-636 additions alone (top 12):

VariableFamilyDropSE
abilitiesOwnrules0.340.11
ppmDiscountsize0.160.16
carriedDistinctloadout0.110.07
Ld_bestsize0.080.05
ptsAtMinsize0.060.05
ptsAtMaxsize0.030.12
fx_partlyrules0.030.08
sizeMinsize0.030.05
modelProfilessize0.030.03
ppmAtMinsize0.030.03
carriedTotalloadout0.010.03
dsRangedloadout0.010.01
FamilyDropSE
ledger12.191.20
killclass2.340.43
size0.810.24
interaction0.380.35
datasheet0.170.27
weaponkw0.000.00
keywords-0.030.03
weapons-0.220.47
flags-0.290.23
rules-0.490.29
loadout-0.620.21
all T-636 additions-0.660.39

Permutation importance, held out, model a1 (drop in pairwise vs consensus, points):

VariableDropSE
c_kill19.870.99
c_presence1.530.40
c_hold0.610.21
c_soak0.590.23
c_score0.460.21
c_spawn0.090.07
M0.000.00
T0.000.00
Sv0.000.00
inv0.000.00
W0.000.00
Ld0.000.00
OC0.000.00
models0.000.00
pts0.000.00

The T-636 additions alone (top 12):

VariableFamilyDropSE
kwu_Vehiclekeywords0.000.00
kwu_Walkerkeywords0.000.00
kwu_Mountedkeywords0.000.00
kwu_Aircraftkeywords0.000.00
kwu_Transportkeywords0.000.00
kwu_Dedicated_Transportkeywords0.000.00
kwu_Leaderkeywords0.000.00
kwu_Grenadeskeywords0.000.00
kwu_Smokekeywords0.000.00
kwu_Framekeywords0.000.00
kwu_Tacticuskeywords0.000.00
kwu_Daemonkeywords0.000.00
FamilyDropSE
ledger21.211.13
datasheet0.000.00
weapons0.000.00
flags0.000.00
killclass0.000.00
interaction0.000.00
keywords0.000.00
size0.000.00
loadout0.000.00
weaponkw0.000.00
rules0.000.00
all T-636 additions0.000.00

Partial dependence, model c (fit on all 52; consensus scale; biggest step):

Partial dependence, model d (fit on all 52; consensus scale; biggest step):

The trees' splits (fit on all 52): variable, gain, median threshold, uses:

The trees' interactions (a depth-3 fit on all 52, 100 trees: one split under the other, by the child's gain):

Linear (b) fitted on all 52, largest standardised weights:

The ledger's seven weights re-fitted on all 52 (a1):

Units the best model (c) gets right that the ledger gets wrong (held-out share of the unit's pairs ordered right):

UnitConsensusConsensus rankLedger rankBest modelLedgerGainTop variables (percentile)
Harpy0.850.533100.0%65.6%34kw_is_burrowers 0, fx_arrival 0, inv 25, kw_precision 0, kw_is_fly 76
Toxicrene1.33346.52386.9%53.4%33kw_is_burrowers 0, fx_arrival 0, inv 25, kw_precision 0, kw_is_fly 0
Gargoyles3.6194764.9%41.2%24kw_is_burrowers 0, fx_arrival 0, inv 25, kw_precision 0, kw_is_fly 76
Tyranid Warriors with Ranged Bio-Weapons1.8431362.8%44.0%19kw_is_burrowers 0, fx_arrival 0, inv 25, kw_precision 0, kw_is_fly 0
Tyrant Guard2.832.54981.5%65.4%16kw_is_burrowers 0, fx_arrival 0, inv 25, kw_precision 0, kw_is_fly 0
Genestealers3.8143468.6%55.3%13kw_is_burrowers 0, fx_arrival 0, inv 20, kw_precision 0, kw_is_fly 0
Neurotyrant3.667173965.0%52.3%13kw_is_burrowers 0, fx_arrival 0, inv 0, kw_precision 0, kw_is_fly 76
Norn Emissary4.2111974.0%61.9%12kw_is_burrowers 0, fx_arrival 0, inv 0, kw_precision 90, kw_is_fly 0
Ripper Swarms2403793.5%83.2%10kw_is_burrowers 0, fx_arrival 0, inv 25, kw_precision 0, kw_is_fly 0
Carnifexes2403081.0%71.9%9kw_is_burrowers 0, fx_arrival 0, inv 25, kw_precision 0, kw_is_fly 0

And the reverse (the ledger right, the model wrong):

Every flexible model misses the same way (mean held-out prediction minus consensus, letters):