Tactical Reroll
⋯

T-632 tables (tools/diagnose.js output)

From reports/tiers/2026-10-04-variables-diagnosis-tables.md , rendered when the site is built.

Ceiling (each list against the mean of the other four): 88.5%

ModelIn-sample vs consensus (tuned)In-sample, most flexibleIn-sample vs listsHeld out vs consensus (± SE)Held out vs lists (± SE)Held out, pooled across folds (consensus / lists)Scrambled vs consensus (SD)Scrambled vs lists (SD)
(a) the ledger as it is (step 6)73.3%73.3%74.5%72.9% ± 0.474.1% ± 0.670.2% / 71.4%50.8% (4.5)51.7% (3.7)
(a1) the ledger's form, its seven weights re-fitted73.3%73.8%74.5%69.7% ± 1.070.1% ± 1.066.8% / 67.5%49.1% (3.8)50.7% (3.9)
(b) linear on every variable (ridge or lasso)99.1%99.1%92.6%70.0% ± 1.068.6% ± 1.169.1% / 67.8%45.8% (8.0)46.4% (8.2)
(b2) linear on a short core list (25 variables), ridge72.2%81.7%71.0%66.0% ± 0.964.2% ± 0.864.1% / 62.4%46.5% (8.4)47.1% (9.2)
(c) linear plus every pairwise interaction (kernel, ridge)96.5%96.5%92.1%72.6% ± 0.770.6% ± 0.970.8% / 69.5%43.0% (7.0)43.9% (7.3)
(d) boosted trees, monotone where the sign is obvious96.0%99.8%91.8%67.2% ± 0.866.7% ± 0.865.9% / 66.0%50.2% (4.9)50.8% (4.8)
(e) nearest neighbours (similar units)100.0%100.0%92.4%64.6% ± 1.462.0% ± 1.461.6% / 59.6%43.6% (8.5)44.6% (9.1)

Learning curve (trained on m units, scored on the pairs among the rest; ± SE over splits):

Per list, held out:

Modelauspexhivemindsecondmaelstromastrategas
ceiling91.7%87.7%81.9%94.9%86.5%
a79.9%78.4%76.2%70.8%65.1%
a176.7%74.3%71.9%66.4%61.3%
b75.8%67.7%71.0%65.3%63.4%
b273.2%65.0%63.8%58.3%60.6%
c77.4%70.5%72.5%68.7%64.1%
d75.2%68.1%72.4%60.4%57.6%
e69.2%61.9%64.0%58.7%56.1%

Settings chosen in the folds (most often):

Flexible against the ledger, held out (paired; SE over repeats; bootstrap over units SD and the share of resamples at or below 0):

Pairvs consensusvs listsbootstrap SDP(≤ 0)scrambled SD
a1 − a-3.2 ± 0.9-4.0 ± 1.02.20.992.3
b − a-2.9 ± 0.8-5.4 ± 1.04.60.7210.9
b2 − a-7.0 ± 0.9-9.9 ± 0.85.80.8910.2
c − a-0.4 ± 0.4-3.4 ± 0.64.70.5310.2
d − a-5.7 ± 0.5-7.3 ± 0.74.10.926.8
e − a-8.3 ± 1.3-12.1 ± 1.44.90.9511.5
b − a10.3 ± 1.1-1.5 ± 1.33.60.489.9
b2 − a1-3.7 ± 1.5-5.9 ± 1.55.00.769.1
c − a12.9 ± 0.90.5 ± 1.13.50.239.4
d − a1-2.5 ± 0.8-3.4 ± 1.03.10.796.6
e − a1-5.1 ± 1.5-8.1 ± 1.44.10.8811.1

Permutation importance, held out, model c (drop in pairwise vs consensus, points):

VariableDropSE
kw_precision0.960.13
arrives0.800.32
kw_psychic0.730.14
inv0.660.16
value0.660.22
kw_is_battleline0.620.24
fx_arrival0.610.20
fx_cp0.610.13
ix_invxT0.600.20
kw_is_fly0.550.18
kw_is_burrowers0.530.14
valueCheap0.490.17
k6_chaff0.490.16
lands0.480.17
lr_movement0.480.15

The T-636 additions alone (top 12):

VariableFamilyDropSE
modelProfilessize0.420.12
OC_maxsize0.380.15
scouts_inchesrules0.320.07
anti_infantryweaponkw0.280.16
fx_partlyrules0.270.15
kwu_Aircraftkeywords0.230.10
transportCapacityrules0.210.10
fx_condEffectsrules0.180.15
wk_twinLinkedAnyweaponkw0.180.12
formSharesize0.170.10
ppmDiscountsize0.170.14
maxAP_anyloadout0.140.17
FamilyDropSE
flags4.420.64
weapons2.960.34
ledger2.740.54
all T-636 additions2.070.46
killclass0.980.23
rules0.960.32
datasheet0.710.30
keywords0.410.16
size0.390.27
interaction0.140.23
weaponkw0.020.20
loadout-0.480.33

Permutation importance, held out, model d (drop in pairwise vs consensus, points):

VariableDropSE
c_score4.120.75
value2.600.49
ix_arrivexMelee1.010.34
k6_monster0.670.30
ppmDiscount0.660.21
inv0.590.16
k6_light_vehicle0.530.25
valueCheap0.490.24
kw_heavy0.420.27
lands0.360.24
kw_precision0.350.17
c_kill0.230.11
ptsAtMax0.220.23
c_hold0.200.11
OC0.170.10

The T-636 additions alone (top 12):

VariableFamilyDropSE
ppmDiscountsize0.660.21
ptsAtMaxsize0.220.23
abilitiesOwnrules0.170.08
shortRangedloadout0.100.08
dsMeleeloadout0.070.08
rulesCountrules0.060.10
sizeRatiosize0.060.03
W_maxsize0.040.02
kwu_Transportkeywords0.030.03
fx_partlyrules0.030.07
defaultWeaponsloadout0.010.01
maxS_anyloadout0.010.01
FamilyDropSE
ledger10.391.00
size0.970.35
killclass0.630.45
interaction0.540.45
datasheet0.450.27
keywords0.000.00
weaponkw-0.010.01
weapons-0.060.46
rules-0.160.21
loadout-0.190.28
flags-0.250.19
all T-636 additions-0.340.33

Permutation importance, held out, model a1 (drop in pairwise vs consensus, points):

VariableDropSE
c_kill17.661.20
c_score1.400.30
c_presence1.060.43
c_hold0.980.24
c_soak0.850.24
c_spawn0.180.06
c_actions0.100.04
M0.000.00
T0.000.00
Sv0.000.00
inv0.000.00
W0.000.00
Ld0.000.00
OC0.000.00
models0.000.00

The T-636 additions alone (top 12):

VariableFamilyDropSE
kwu_Vehiclekeywords0.000.00
kwu_Walkerkeywords0.000.00
kwu_Mountedkeywords0.000.00
kwu_Aircraftkeywords0.000.00
kwu_Transportkeywords0.000.00
kwu_Dedicated_Transportkeywords0.000.00
kwu_Leaderkeywords0.000.00
kwu_Grenadeskeywords0.000.00
kwu_Smokekeywords0.000.00
kwu_Framekeywords0.000.00
kwu_Tacticuskeywords0.000.00
kwu_Daemonkeywords0.000.00
FamilyDropSE
ledger19.041.03
datasheet0.000.00
weapons0.000.00
flags0.000.00
killclass0.000.00
interaction0.000.00
keywords0.000.00
size0.000.00
loadout0.000.00
weaponkw0.000.00
rules0.000.00
all T-636 additions0.000.00

Partial dependence, model c (fit on all 52; consensus scale; biggest step):

Partial dependence, model d (fit on all 52; consensus scale; biggest step):

The trees' splits (fit on all 52): variable, gain, median threshold, uses:

The trees' interactions (a depth-3 fit on all 52, 100 trees: one split under the other, by the child's gain):

Linear (b) fitted on all 52, largest standardised weights:

The ledger's seven weights re-fitted on all 52 (a1):

Units the best model (c) gets right that the ledger gets wrong (held-out share of the unit's pairs ordered right):

UnitConsensusConsensus rankLedger rankBest modelLedgerGainTop variables (percentile)
Toxicrene1482386.2%50.7%36kw_precision 0, arrives 27, kw_psychic 0, inv 25, value 57
Harpy0.66751.53393.6%61.3%32kw_precision 0, arrives 0, kw_psychic 0, inv 25, value 37
Gargoyles3.75154765.5%35.2%30kw_precision 0, arrives 27, kw_psychic 0, inv 25, value 10
Tyranid Warriors with Ranged Bio-Weapons240.51369.7%48.9%21kw_precision 0, arrives 27, kw_psychic 0, inv 25, value 76
Tyranid Prime with Lash Whip2.5351469.9%51.3%19kw_precision 0, arrives 0, kw_psychic 0, inv 25, value 75
Genestealers3.75153471.2%53.6%18kw_precision 0, arrives 27, kw_psychic 0, inv 20, value 35
Norn Emissary4.2111975.6%61.2%14kw_precision 90, arrives 27, kw_psychic 100, inv 0, value 65
Tyrant Guard2.829.54974.8%60.4%14kw_precision 0, arrives 27, kw_psychic 0, inv 25, value 6
Neurotyrant3.667173961.2%47.2%14kw_precision 0, arrives 27, kw_psychic 92, inv 0, value 25
Carnifexes240.53086.2%74.7%12kw_precision 0, arrives 0, kw_psychic 0, inv 25, value 43

And the reverse (the ledger right, the model wrong):

Every flexible model misses the same way (mean held-out prediction minus consensus, letters):