Tactical Reroll
⋯

T-632 tables (tools/diagnose.js output)

From reports/tiers/2026-10-03-diagnosis-tables.md , rendered when the site is built.

Ceiling (each list against the mean of the other four): 88.5%

ModelIn-sample vs consensus (tuned)In-sample, most flexibleIn-sample vs listsHeld out vs consensus (± SE)Held out vs lists (± SE)Held out, pooled across folds (consensus / lists)Scrambled vs consensus (SD)Scrambled vs lists (SD)
(a) the ledger as it is (step 6)73.3%73.3%74.5%72.9% ± 0.474.1% ± 0.670.2% / 71.4%50.8% (4.5)51.7% (3.7)
(a1) the ledger's form, its seven weights re-fitted73.3%73.8%74.5%69.7% ± 1.070.1% ± 1.066.8% / 67.5%49.1% (3.8)50.7% (3.9)
(b) linear on every variable (ridge or lasso)89.1%98.8%87.7%69.7% ± 1.368.6% ± 0.969.2% / 68.2%45.9% (6.9)46.5% (7.5)
(b2) linear on a short core list (25 variables), ridge72.2%81.7%71.0%66.0% ± 0.964.2% ± 0.864.1% / 62.4%46.5% (8.4)47.1% (9.2)
(c) linear plus every pairwise interaction (kernel, ridge)97.2%97.2%92.2%73.3% ± 0.971.4% ± 1.071.5% / 70.6%44.8% (6.6)45.8% (7.0)
(d) boosted trees, monotone where the sign is obvious95.3%99.7%91.5%68.3% ± 0.967.6% ± 0.866.6% / 66.5%49.3% (3.8)50.1% (4.5)
(e) nearest neighbours (similar units)100.0%100.0%92.3%71.9% ± 0.870.1% ± 1.169.7% / 68.4%45.2% (7.5)46.3% (7.5)

Learning curve (trained on m units, scored on the pairs among the rest; ± SE over splits):

Per list, held out:

Modelauspexhivemindsecondmaelstromastrategas
ceiling91.7%87.7%81.9%94.9%86.5%
a79.9%78.4%76.2%70.8%65.1%
a176.7%74.3%71.9%66.4%61.3%
b75.4%68.1%69.5%65.3%64.8%
b273.2%65.0%63.8%58.3%60.6%
c77.9%71.7%71.2%69.9%66.5%
d76.0%69.0%72.5%61.4%59.1%
e76.6%71.2%69.1%69.4%64.2%

Settings chosen in the folds (most often):

Flexible against the ledger, held out (paired; SE over repeats; bootstrap over units SD and the share of resamples at or below 0):

Pairvs consensusvs listsbootstrap SDP(≤ 0)scrambled SD
a1 − a-3.2 ± 0.9-4.0 ± 1.02.20.992.3
b − a-3.2 ± 1.1-5.5 ± 1.04.30.749.9
b2 − a-7.0 ± 0.9-9.9 ± 0.85.80.8910.2
c − a0.3 ± 0.7-2.6 ± 0.74.70.479.5
d − a-4.6 ± 0.7-6.5 ± 0.64.00.916.1
e − a-1.1 ± 0.7-4.0 ± 0.84.80.6311.1
b − a10.0 ± 1.1-1.5 ± 1.13.40.548.6
b2 − a1-3.7 ± 1.5-5.9 ± 1.55.00.769.1
c − a13.6 ± 0.81.3 ± 0.93.60.178.7
d − a1-1.4 ± 1.2-2.5 ± 1.33.00.705.9
e − a12.2 ± 1.3-0.0 ± 1.43.90.3610.4

Permutation importance, held out, model c (drop in pairwise vs consensus, points):

VariableDropSE
kw_is_battleline1.630.21
kw_is_transport1.140.12
kw_precision1.130.15
kw_is_burrowers0.850.18
k6_light_vehicle0.830.21
inv0.810.15
arrives0.810.22
kw_psychic0.790.15
kw_heavy0.710.17
soakFight0.660.15
valueCheap0.660.19
lands0.600.19
bestR_A0.590.12
kw_is_fly0.570.24
rangedAttacks1000.430.13
FamilyDropSE
flags6.260.72
weapons3.570.53
ledger3.410.68
killclass1.500.31
datasheet0.980.37
interaction0.570.23

Permutation importance, held out, model d (drop in pairwise vs consensus, points):

VariableDropSE
c_score4.700.46
value1.800.42
ix_arrivexMelee1.660.39
k6_monster1.210.26
valueCheap1.140.36
lands1.110.29
k6_light_vehicle0.880.20
kw_precision0.510.22
ix_MxMelee0.450.24
kw_heavy0.350.20
kw_psychic0.250.10
k6_character0.230.10
kw_ignoresCover0.200.07
ocPer1000.180.10
c_kill0.150.16
FamilyDropSE
ledger12.411.05
interaction1.920.47
killclass0.740.45
datasheet0.270.32
weapons0.180.28
flags-0.180.14

Permutation importance, held out, model a1 (drop in pairwise vs consensus, points):

VariableDropSE
c_kill17.741.18
c_score1.810.25
c_presence1.650.53
c_hold1.080.22
c_soak0.390.24
c_spawn0.100.07
c_actions0.040.04
M0.000.00
T0.000.00
Sv0.000.00
inv0.000.00
W0.000.00
Ld0.000.00
OC0.000.00
models0.000.00
FamilyDropSE
ledger20.501.38
datasheet0.000.00
weapons0.000.00
flags0.000.00
killclass0.000.00
interaction0.000.00

Partial dependence, model c (fit on all 52; consensus scale; biggest step):

Partial dependence, model d (fit on all 52; consensus scale; biggest step):

The trees' splits (fit on all 52): variable, gain, median threshold, uses:

The trees' interactions (a depth-3 fit on all 52, 100 trees: one split under the other, by the child's gain):

Linear (b) fitted on all 52, largest standardised weights:

The ledger's seven weights re-fitted on all 52 (a1):

Units the best model (c) gets right that the ledger gets wrong (held-out share of the unit's pairs ordered right):

UnitConsensusConsensus rankLedger rankBest modelLedgerGainTop variables (percentile)
Toxicrene1482389.2%50.7%39kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 33
Harpy0.66751.53393.6%61.3%32kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 29
Gargoyles3.75154765.6%35.2%30kw_is_battleline 96, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 24
Tyranid Prime with Lash Whip2.5351471.0%51.3%20kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 82
Tyrant Guard2.829.54978.0%60.4%18kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 37
Genestealers3.75153470.8%53.6%17kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 49
Neurotyrant3.667173964.0%47.2%17kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 12
Norn Emissary4.2111975.6%61.2%14kw_is_battleline 0, kw_is_transport 0, kw_precision 90, kw_is_burrowers 0, k6_light_vehicle 51
Carnifexes240.53089.2%74.7%14kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 96
Tyranid Warriors with Ranged Bio-Weapons240.51361.5%48.9%13kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 18

And the reverse (the ledger right, the model wrong):

Every flexible model misses the same way (mean held-out prediction minus consensus, letters):